Thuta Learning
Computer Vision
AdvancedAIintermediate

Object Detection Architecture များ

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • Object Detection Architecture များ concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • နမူနာ code ကို ကိုယ်တိုင် run ပြီး output စစ်နိုင်ရန်
  • Tutorial Platform project နှင့် production scenario တွင် မှန်ကန်စွာအသုံးချနိုင်ရန်

နားလည်ထားရမယ့် အချက်

R-CNN လိုမျိုး ရှေးဦး object detector တွေဟာ separate stage နှစ်ခုနဲ့ အလုပ်လုပ်ခဲ့ကြပါတယ်— ပထမဆုံး region-proposal အဆင့်က ပုံကို scan လုပ်ပြီး object ပါနိုင်တဲ့ candidate box ထောင်ချီကို အကြံပြုပါတယ်၊ ပြီးရင် ဒုတိယ network တစ်ခုက candidate တစ်ခုချင်းစီကို crop လုပ်ပြီး classify လုပ်ပါတယ်။ ဒါဟာ accuracy ကောင်းပေမယ့် နှေးပါတယ်၊ classification network ကို proposed region တစ်ခုစီအတွက် တစ်ကြိမ်စီ run ရလို့ပါ။ YOLO ('You Only Look Once') နဲ့ တခြား single-shot detector တွေကတော့ လုံးဝ ကွဲပြားတဲ့ ချဉ်းကပ်မှုကို သုံးပါတယ်— proposal အဆင့်ကို လုံးဝ ကျော်ပြီး၊ ပုံကို grid ပုံသေတစ်ခုအဖြစ် ခွဲကာ network ကို forward pass တစ်ခါတည်းနဲ့ grid cell တိုင်းရဲ့ အထဲမှာ ဘာရှိသလဲဆိုတာ တစ်ပြိုင်နက် predict ခိုင်းပါတယ်။ cell တစ်ခုစီဟာ ကိုယ့် အထဲမှာ center ကျရောက်နေတဲ့ object ကို ဖော်ထုတ်ဖို့ တာဝန်ရှိပြီး— network က region တွေကို အစဉ်လိုက် 'ကြည့်' တာမျိုး မဟုတ်ဘဲ shared feature map တစ်ခုတည်းကနေ prediction အားလုံးကို တစ်ပြိုင်နက် ထုတ်ပေးပါတယ်။

တိတိကျကျ ပြောရရင် single-shot detector တစ်ခုရဲ့ grid cell တစ်ခုစီက output ဟာ 'anchor' (pedestrian အတွက် ရှည်ချောင်တဲ့ box၊ car အတွက် ကျယ်ပြန့်တဲ့ box လိုမျိုး ကြိုတင်သတ်မှတ်ထားတဲ့ box ပုံသဏ္ဌာန်) တစ်ခုစီအတွက် ထပ်ခါထပ်ခါ ရှိနေတဲ့ fixed-size vector တစ်ခုပါ။ anchor တစ်ခုစီအတွက် network က objectness score (ဒီနေရာမှာ object တစ်ခုခု ရှိသလား?) တစ်ခုနဲ့ box-coordinate offset 4 ခု (anchor ရဲ့ default position နဲ့ size ကို actual object နဲ့ ကိုက်ညီအောင် ချိန်ညှိချက်) ကို predict လုပ်ပါတယ်။ ဒါကြောင့် anchor 3 ခုနဲ့ value 5 ခုစီ ရှိတဲ့ grid cell တစ်ခုဟာ spatial location တစ်ခုမှာ output channel 15 ခု လိုအပ်ပါတယ်။ ဒါကြောင့် detection head ဟာ backbone ရဲ့ feature map အပေါ်မှာ ထပ်တင်ထားတဲ့ convolutional layer သေးငယ်တစ်ခုပါပဲ— in_channels ကို num_anchors * num_outputs အဖြစ် map လုပ်တဲ့ 1x1 (ဒါမှမဟုတ် 3x3) convolution တစ်ခုက grid cell တစ်ခုစီအတွက် prediction bundle တစ်ခုကို dense tensor တစ်ခုတည်းအဖြစ် ထုတ်ပေးပါတယ်— crop လုပ်စရာ၊ region တစ်ခုစီအတွက် network ခေါ်ဆိုစရာ မလိုအပ်တော့ဘဲ၊ ဒါက single-shot detection ကို real-time သုံးဖို့ လုံလောက်စွာ မြန်ဆန်စေတဲ့ အကြောင်းရင်းအတိအကျပါပဲ။

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

Tutorial Platform မှာ lesson content ထဲက embed လုပ်ထားတဲ့ screenshot များအတွင်းရှိ UI element များ — button, code block, input field များ — ကို အလိုအလျောက် ရှာဖွေတည်နေရာသတ်မှတ်ဖို့ ဒီလိုမျိုး single-shot detection head ကို သုံးနိုင်ပါတယ်၊ bounding box တွေကို ထုတ်ပေးပြီး auto-generated alt-text ('Submit ခလုတ်အစိမ်းရောင် ပြသနေတဲ့ screenshot') နဲ့ interactive click-to-zoom hotspot နှစ်ခုစလုံးကို ဖန်တီးပေးနိုင်ပါတယ်၊ author က lesson အသစ်တစ်ခု upload လုပ်တိုင်း screenshot တွေကို အစုလိုက် process လုပ်ရမှာမို့ ပိုနှေးတဲ့ two-stage pipeline မလိုအပ်တော့ပါဘူး။

အတူတူ စမ်းရေးကြည့်မယ်

python
import torch
import torch.nn as nn

class TinyDetectionHead(nn.Module):
    def __init__(self, in_channels, num_anchors=3, num_outputs=5):
        super().__init__()
        self.backbone = nn.Sequential(
            nn.Conv2d(3, 16, kernel_size=3, stride=2, padding=1),          # 32x32 -> 16x16
            nn.ReLU(),
            nn.Conv2d(16, in_channels, kernel_size=3, stride=2, padding=1),  # 16x16 -> 8x8
            nn.ReLU(),
        )
        self.head = nn.Conv2d(in_channels, num_anchors * num_outputs, kernel_size=1)
        self.num_anchors = num_anchors
        self.num_outputs = num_outputs

    def forward(self, x):
        features = self.backbone(x)                      # (batch, in_channels, grid_h, grid_w)
        raw = self.head(features)                         # (batch, num_anchors*num_outputs, grid_h, grid_w)
        batch, _, grid_h, grid_w = raw.shape
        return raw.view(batch, self.num_anchors, self.num_outputs, grid_h, grid_w)

torch.manual_seed(0)
model = TinyDetectionHead(in_channels=32)
images = torch.randn(4, 3, 32, 32)  # fake batch of 4 RGB images
predictions = model(images)
print(predictions.shape)
You should see
torch.Size([4, 3, 5, 8, 8]) — backbone ရဲ့ stride-2 convolution နှစ်ခုက 32x32 input ကို 8x8 feature grid အဖြစ် ချုံ့ပေးပြီး၊ detection head က grid cell 64 ခုစလုံးအတွက် anchor 3 ခု x value 5 ခုစီ (objectness score 1 ခု + box coordinate 4 ခု) ကို ထုတ်ပေးပါတယ်၊ ဒါကို (batch, num_anchors, num_outputs, grid_h, grid_w) tensor အဖြစ် reshape လုပ်ထားပါတယ်။

၅ မိနစ် စမ်းကြည့်

num_anchors ကို 3 ကနေ 5 ကို ပြောင်းပြီး model ကို ပြန် run ပါ၊ ပြီးရင် self.head ကနေ ထွက်လာတဲ့ (reshape မလုပ်ခင်) raw channel count ဟာ num_anchors * num_outputs (25) နဲ့ညီသလား၊ နောက်ဆုံး print ထုတ်တဲ့ shape ရဲ့ ဒုတိယ dimension က 5 ကို update ဖြစ်သွားသလား ကိုယ်တိုင် တွက်ချက် စစ်ဆေးကြည့်ပါ— ဒါဟာ တကယ့် detector တွေက cell တစ်ခုစီအတွက် anchor shape ပိုများခြင်းကို output tensor ပိုကြီးလာခြင်းနဲ့ ဘယ်လို လဲလှယ်ကြောင်းကို ထင်ဟပ်ပါတယ်။

သတိလေးတစ်ချက်

နောက်ဆုံး .view() ခေါ်ဆိုမှုမှာ grid_h နဲ့ grid_w ကို raw.shape ကနေ ဖတ်မယ့်အစား hard-code လုပ်ထားရင်၊ စမ်းသပ်ခဲ့တာနဲ့ မတူတဲ့ input resolution တစ်ခု ပေးလိုက်တာနဲ့ model ပျက်သွားပါလိမ့်မယ်။

raw objectness value ကို sigmoid အရင်မသုံးဘဲ probability တစ်ခုအဖြစ် သတ်မှတ်လိုက်ရင် [0, 1] အပြင်ဘက်က အဓိပ္ပာယ်မရှိတဲ့ score တွေ ထွက်လာပါလိမ့်မယ်၊ head ရဲ့ နောက်ဆုံး conv မှာ ကိုယ်ပိုင် activation function မရှိလို့ပါပဲ။

Wikipedia — You Only Look OnceComputer Vision

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • နောက်ဆုံး .view() ခေါ်ဆိုမှုမှာ grid_h နဲ့ grid_w ကို raw.shape ကနေ ဖတ်မယ့်အစား hard-code လုပ်ထားရင်၊ စမ်းသပ်ခဲ့တာနဲ့ မတူတဲ့ input resolution တစ်ခု ပေးလိုက်တာနဲ့ model ပျက်သွားပါလိမ့်မယ်။
  • raw objectness value ကို sigmoid အရင်မသုံးဘဲ probability တစ်ခုအဖြစ် သတ်မှတ်လိုက်ရင် [0, 1] အပြင်ဘက်က အဓိပ္ပာယ်မရှိတဲ့ score တွေ ထွက်လာပါလိမ့်မယ်၊ head ရဲ့ နောက်ဆုံး conv မှာ ကိုယ်ပိုင် activation function မရှိလို့ပါပဲ။
  • နမူနာ code ကို production system ပေါ် တိုက်ရိုက်မစမ်းဘဲ local/test environment တွင် အရင်အတည်ပြုပါ။

လေ့ကျင့်ခန်း

num_anchors ကို 3 ကနေ 5 ကို ပြောင်းပြီး model ကို ပြန် run ပါ၊ ပြီးရင် self.head ကနေ ထွက်လာတဲ့ (reshape မလုပ်ခင်) raw channel count ဟာ num_anchors * num_outputs (25) နဲ့ညီသလား၊ နောက်ဆုံး print ထုတ်တဲ့ shape ရဲ့ ဒုတိယ dimension က 5 ကို update ဖြစ်သွားသလား ကိုယ်တိုင် တွက်ချက် စစ်ဆေးကြည့်ပါ— ဒါဟာ တကယ့် detector တွေက cell တစ်ခုစီအတွက် anchor shape ပိုများခြင်းကို output tensor ပိုကြီးလာခြင်းနဲ့ ဘယ်လို လဲလှယ်ကြောင်းကို ထင်ဟပ်ပါတယ်။

You'll know it worked when: torch.Size([4, 3, 5, 8, 8]) — backbone ရဲ့ stride-2 convolution နှစ်ခုက 32x32 input ကို 8x8 feature grid အဖြစ် ချုံ့ပေးပြီး၊ detection head က grid cell 64 ခုစလုံးအတွက် anchor 3 ခု x value 5 ခုစီ (objectness score 1 ခု + box coordinate 4 ခု) ကို ထုတ်ပေးပါတယ်၊ ဒါကို (batch, num_anchors, num_outputs, grid_h, grid_w) tensor အဖြစ် reshape လုပ်ထားပါတယ်။

Object Detection Architecture များ | Thuta Learning