Thuta Learning
Computer Vision
BasicAIintermediate

Classical Feature Detection

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • Classical Feature Detection concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • နမူနာ code ကို ကိုယ်တိုင် run ပြီး output စစ်နိုင်ရန်
  • Tutorial Platform project နှင့် production scenario တွင် မှန်ကန်စွာအသုံးချနိုင်ရန်

နားလည်ထားရမယ့် အချက်

keypoint ကောင်းတစ်ခု — Harris corner (သို့) SIFT ကဲ့သို့ classical detector များ ရှာဖွေနေသော အမျိုးအစား — ဆိုသည်မှာ local image content သည် ယုံကြည်စိတ်ချစွာ ပြန်လည်သတ်မှတ်နိုင်လောက်အောင် ထူးခြားသော နေရာတစ်ခု ဖြစ်ပြီး၊ 'ထူးခြားခြင်း' ကို formalize လုပ်သည့် standard နည်းလမ်းမှာ window သေးငယ်တစ်ခုကို လမ်းကြောင်းမည်သည့်ဘက်ကိုမဆို အနည်းငယ် ရွှေ့လိုက်လျှင် pixel value များ ဘာဖြစ်လာသလဲဟု မေးခြင်းပင် ဖြစ်သည်။ တပြေးညီ ညီညာသော ဧရိယာတစ်ခုတွင် window ရွှေ့လိုက်ခြင်းသည် ဘာမျှ သိပ်မပြောင်းလဲပါ — ထို့နေရာတွင် ထူးခြားသည့်အရာ မရှိပါ။ ဖြောင့်တန်းသော edge တစ်လျှောက်တွင်လည်း edge နှင့်ပြိုင်၍ window ရွှေ့လိုက်ခြင်းသည် edge ကို ဖြတ်၍ ရွှေ့သည်ထက် ဘာမျှ သိပ်မပြောင်းလဲပါ — edge တစ်ခုသည် ဦးတည်ချက်တစ်ခုတည်းတွင်သာ ထူးခြားခြင်းကြောင့် gradient ပြင်းထန်သော်လည်း keypoint ကောင်းတစ်ခု မဖြစ်တတ်ပါ။ corner အစစ်တစ်ခုတွင်တော့ window ကို မည်သည့်ဘက် ရွှေ့ရွှေ့ content သည် များစွာ ပြောင်းလဲသွားသည် — အကြောင်းမှာ edge နှစ်ခု ထိုနေရာတွင် ဆုံသည့်အတွက် ဖြစ်သည် — ဤ ဦးတည်ချက်နှစ်ခု property ကို Harris ကဲ့သို့ corner detector အစစ်တစ်ခုက (structure tensor မှတဆင့်) တိတိကျကျ စစ်ဆေးပေးပြီး ဒီသင်ခန်းစာ code တွက်ချက်သော ရိုးရှင်းသော ဦးတည်ချက်တစ်ခုတည်း gradient magnitude ထက် တစ်ဆင့်ပို၍ လုပ်ဆောင်ပေးသည်။

SIFT, Harris corner ကဲ့သို့ classical feature detector များသည် 'ထူးခြားပြီး ပြန်ရှာတွေ့နိုင်ခြင်း' ဆိုသည်မှာ ဘာလဲဆိုသည့် လူ့ hypothesis တစ်ခုကို မည်သူမဆို ဖတ်ပြီး စဉ်းစားနိုင်၊ training data မလိုဘဲ run နိုင်သော ရှင်းလင်းသော mathematical test တစ်ခုအဖြစ် ဖော်ပြထားသည်။ CNN တစ်ခုတွင်တော့ ထိုကဲ့သို့ ရှင်းလင်းသော hypothesis မရှိပါ — ၎င်း architecture အတွင်း 'corner ကို ရှာဖွေပါ' ဟု ဘယ်နေရာမှ ညွှန်ကြားထားခြင်း မရှိပါ။ သို့သော် trained CNN တစ်ခု၏ early convolutional filter များ အမှန်တကယ် ဘာကို respond လုပ်သလဲဟု visualize လုပ်ကြည့်လျှင် လက်ဖြင့်ဒီဇိုင်းဆွဲထားသော ဟာများနှင့် အံ့သြစရာကောင်းလောက်အောင် ဆင်တူသော edge, blob detector များကို မကြာခဏ တွေ့ရသည် — ထိုပုံစံများသည် training loss ကို လျှော့ချရာတွင် အသုံးဝင်ကြောင်း ပေါ်ပေါက်လာသောကြောင့် သန့်စင်စွာ သင်ယူထားခြင်းသာ ဖြစ်သည်။ deeper layer များကတော့ ပိုမို ရှေ့ဆက်သွားပြီး pattern ပေါင်းစပ်မှုများ — အမွှေးအမှင် texture, ဘီးပုံသဏ္ဍာန် curve, မျက်လုံးပုံ blob — အတွက် detector များကို သင်ယူသည် — ၎င်းတို့အတွက် mathematical rule တစ်ခုကို လူတစ်ဦးက လက်ဖြင့်ဒီဇိုင်းဆွဲရန် အလွန်ခက်ခဲမည် ဖြစ်သည်။ ထို့ကြောင့်ပင် image stitching (သို့) simple tracking ကဲ့သို့ ဈေးသက်သာပြီး interpretable ဖြစ်သော၊ training မလိုသော task များအတွက် classical detector များ ယနေ့တိုင် အသုံးဝင်နေဆဲဖြစ်ပြီး၊ interest ရှိသော pattern သည် လက်ဖြင့်ဖော်ပြရန် ရှုပ်ထွေးလွန်းသည့်နေရာတိုင်းတွင် learned feature များက လွှမ်းမိုးထားသည်။

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

Tutorial Platform ၏ course-thumbnail uploader သည် ဈေးကြီးသော learned similarity model ကို ခေါ်မီ ဈေးသက်သာသော classical fingerprint check တစ်ခုကို ဦးစွာ run လုပ်သည် — upload လုပ်လိုက်သော thumbnail အသစ်တိုင်းအတွက် gradient-magnitude 'keypoint density' map တစ်ခု (ဒီသင်ခန်းစာ code ကဲ့သို့ပင်) တွက်ချက်ပြီး ရှိပြီးသား thumbnail များ၏ fingerprint များနှင့် နှိုင်းယှဉ်ကြည့်သည် — instructor တစ်ဦးက banner image အတူတူ (သို့) crop အနည်းငယ်သာ ကွာသည့် ပုံကို ပြန် upload လုပ်လိုက်သည့် ဖြစ်ရပ်များစွာကို ချက်ချင်းနီးပါး ဖမ်းမိစေပြီး၊ ဈေးသက်သာသော classical check က duplicate ဟုတ်၊ မဟုတ် ယုံကြည်စိတ်ချစွာ ဆုံးဖြတ်မရသည့်အခါမှသာ ပိုနှေးသော learned embedding model ကို ပြန်လှည့်သုံးသည်။

အတူတူ စမ်းရေးကြည့်မယ်

python
import torch
import torch.nn.functional as F

torch.manual_seed(0)

# Synthetic grayscale image standing in for a course thumbnail.
image = torch.rand(1, 1, 20, 20)

sobel_x = torch.tensor([[-1., 0., 1.],
                         [-2., 0., 2.],
                         [-1., 0., 1.]]).view(1, 1, 3, 3)
sobel_y = torch.tensor([[-1., -2., -1.],
                         [ 0.,  0.,  0.],
                         [ 1.,  2.,  1.]]).view(1, 1, 3, 3)

grad_x = F.conv2d(image, sobel_x, padding=1)
grad_y = F.conv2d(image, sobel_y, padding=1)
gradient_magnitude = torch.sqrt(grad_x ** 2 + grad_y ** 2)

# A classical corner/keypoint detector is, at its core, looking for
# locations where this gradient magnitude is unusually high in
# *multiple* directions -- flat regions and simple edges score low,
# corner-like regions score high. We approximate that here with a
# blunt threshold instead of a full Harris response.
threshold = gradient_magnitude.mean() + gradient_magnitude.std()
keypoint_mask = gradient_magnitude > threshold

print("gradient_magnitude shape:", gradient_magnitude.shape)
print("keypoint_mask shape:", keypoint_mask.shape)
print("number of candidate keypoints:", keypoint_mask.sum().item())
print("fraction of image flagged as a keypoint:", (keypoint_mask.float().mean() > 0).item())
You should see
gradient_magnitude shape: torch.Size([1, 1, 20, 20]) နှင့် keypoint_mask shape: torch.Size([1, 1, 20, 20]) ကို print ထုတ်သည်။ number of candidate keypoints သည် pixel စုစုပေါင်း 400 ထက် အများကြီး နည်းသော positive integer သေးငယ်တစ်ခု ဖြစ်သည် — threshold (mean + standard deviation တစ်ခု) သည် ဒီဇိုင်းအရ distribution ၏ magnitude အမြင့်ဆုံး tail ကိုသာ ထားခဲ့သဖြင့် pixel အနည်းစုသာ ဖြတ်နိုင်သည် — အတိအကျ count သည် seed ပေါ်မူတည်၍ deterministic ဖြစ်သော်လည်း ဤနေရာတွင် လက်ဖြင့် မတွက်ချက်ထားပါ။ နောက်ဆုံးလိုင်းက fraction of image flagged as a keypoint: True ဟု print ထုတ်သည် — အကြောင်းမှာ constant မဟုတ်သော data ပေါ်တွင် mean + std ၏ အဓိပ္ပာယ်အရ pixel အနည်းဆုံးတစ်ချို့ threshold ထက် ကျော်နေမည် ဖြစ်သောကြောင့်ပင်။

၅ မိနစ် စမ်းကြည့်

blunt mean+std threshold အစား Harris-like response တစ်ခု တွက်ချက်ကြည့်ပါ — Ixx = grad_x**2, Iyy = grad_y**2, Ixy = grad_x*grad_y ကို တွက်ပြီး တစ်ခုစီကို neighborhood သေးငယ်တစ်ခု (ဥပမာ 3x3 average-pooling) ပေါ်မှာ sum လုပ်ကာ response = (Ixx*Iyy - Ixy**2) - 0.04*(Ixx+Iyy)**2 ကို တွက်ပြီး ၎င်းကို threshold လုပ်ကြည့်ပါ — image တစ်ခုတည်းအပေါ် original gradient-magnitude approach နှင့် ယှဉ်လျှင် keypoint ဘယ်နှစ်ခု ကွာခြားစွာ flag လုပ်သလဲ နှိုင်းယှဉ်ကြည့်ပါ။

သတိလေးတစ်ချက်

ဤ demo ကဲ့သို့ raw gradient-magnitude threshold ကို corner detector တစ်ခုဟု ယူဆခြင်း — gradient သည် ဦးတည်ချက်တစ်ခုထက် ပို၍ ကွဲပြားမှု ရှိမရှိ ဘယ်တော့မှ စစ်ဆေးခြင်း မရှိသောကြောင့် ဖြောင့်တန်းသော edge များပေါ်တွင်လည်း corner များအပေါ်ကဲ့သို့ပင် ပြင်းပြင်းထန်ထန် fire ဖြစ်တတ်သည်။

ဤနည်းဖြင့် တွေ့ရှိသော keypoint တစ်ခုသည် scale, rotation invariant ဖြစ်မည်ဟု default အနေနှင့် ယူဆခြင်း — plain gradient-based detection တွင် built-in invariance မရှိပါ — ထို့ကြောင့် thumbnail တူညီသည့် ပုံတစ်ပုံကို resize (သို့) rotate လုပ်လိုက်လျှင် multi-scale (သို့) orientation handling ကို တမင် ထည့်သွင်းထားခြင်း မရှိပါက ယေဘုယျအားဖြင့် ကွဲပြားသော keypoint map ကို ထုတ်ပေးလိမ့်မည် — train ကောင်းစွာ လုပ်ထားသော CNN embedding တစ်ခုကတော့ ထိုကဲ့သို့ ပြောင်းလဲမှုများကို အလိုအလျောက် ပိုမိုကြံ့ခိုင်စွာ ခံနိုင်ရည်ရှိတတ်သည်။

Wikipedia — Feature (computer vision)Computer Vision

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • ဤ demo ကဲ့သို့ raw gradient-magnitude threshold ကို corner detector တစ်ခုဟု ယူဆခြင်း — gradient သည် ဦးတည်ချက်တစ်ခုထက် ပို၍ ကွဲပြားမှု ရှိမရှိ ဘယ်တော့မှ စစ်ဆေးခြင်း မရှိသောကြောင့် ဖြောင့်တန်းသော edge များပေါ်တွင်လည်း corner များအပေါ်ကဲ့သို့ပင် ပြင်းပြင်းထန်ထန် fire ဖြစ်တတ်သည်။
  • ဤနည်းဖြင့် တွေ့ရှိသော keypoint တစ်ခုသည် scale, rotation invariant ဖြစ်မည်ဟု default အနေနှင့် ယူဆခြင်း — plain gradient-based detection တွင် built-in invariance မရှိပါ — ထို့ကြောင့် thumbnail တူညီသည့် ပုံတစ်ပုံကို resize (သို့) rotate လုပ်လိုက်လျှင် multi-scale (သို့) orientation handling ကို တမင် ထည့်သွင်းထားခြင်း မရှိပါက ယေဘုယျအားဖြင့် ကွဲပြားသော keypoint map ကို ထုတ်ပေးလိမ့်မည် — train ကောင်းစွာ လုပ်ထားသော CNN embedding တစ်ခုကတော့ ထိုကဲ့သို့ ပြောင်းလဲမှုများကို အလိုအလျောက် ပိုမိုကြံ့ခိုင်စွာ ခံနိုင်ရည်ရှိတတ်သည်။
  • နမူနာ code ကို production system ပေါ် တိုက်ရိုက်မစမ်းဘဲ local/test environment တွင် အရင်အတည်ပြုပါ။

လေ့ကျင့်ခန်း

blunt mean+std threshold အစား Harris-like response တစ်ခု တွက်ချက်ကြည့်ပါ — Ixx = grad_x**2, Iyy = grad_y**2, Ixy = grad_x*grad_y ကို တွက်ပြီး တစ်ခုစီကို neighborhood သေးငယ်တစ်ခု (ဥပမာ 3x3 average-pooling) ပေါ်မှာ sum လုပ်ကာ response = (Ixx*Iyy - Ixy**2) - 0.04*(Ixx+Iyy)**2 ကို တွက်ပြီး ၎င်းကို threshold လုပ်ကြည့်ပါ — image တစ်ခုတည်းအပေါ် original gradient-magnitude approach နှင့် ယှဉ်လျှင် keypoint ဘယ်နှစ်ခု ကွာခြားစွာ flag လုပ်သလဲ နှိုင်းယှဉ်ကြည့်ပါ။

You'll know it worked when: gradient_magnitude shape: torch.Size([1, 1, 20, 20]) နှင့် keypoint_mask shape: torch.Size([1, 1, 20, 20]) ကို print ထုတ်သည်။ number of candidate keypoints သည် pixel စုစုပေါင်း 400 ထက် အများကြီး နည်းသော positive integer သေးငယ်တစ်ခု ဖြစ်သည် — threshold (mean + standard deviation တစ်ခု) သည် ဒီဇိုင်းအရ distribution ၏ magnitude အမြင့်ဆုံး tail ကိုသာ ထားခဲ့သဖြင့် pixel အနည်းစုသာ ဖြတ်နိုင်သည် — အတိအကျ count သည် seed ပေါ်မူတည်၍ deterministic ဖြစ်သော်လည်း ဤနေရာတွင် လက်ဖြင့် မတွက်ချက်ထားပါ။ နောက်ဆုံးလိုင်းက fraction of image flagged as a keypoint: True ဟု print ထုတ်သည် — အကြောင်းမှာ constant မဟုတ်သော data ပေါ်တွင် mean + std ၏ အဓိပ္ပာယ်အရ pixel အနည်းဆုံးတစ်ချို့ threshold ထက် ကျော်နေမည် ဖြစ်သောကြောင့်ပင်။

Classical Feature Detection | Thuta Learning