Thuta Learning
Computer Vision
BasicAIintermediate

Computer Vision ဆိုတာ ဘာလဲ

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • Computer Vision ဆိုတာ ဘာလဲ concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • နမူနာ code ကို ကိုယ်တိုင် run ပြီး output စစ်နိုင်ရန်
  • Tutorial Platform project နှင့် production scenario တွင် မှန်ကန်စွာအသုံးချနိုင်ရန်

နားလည်ထားရမယ့် အချက်

Computer Vision ဆိုသည်မှာ ကွန်ပျူတာတစ်လုံးအား ဓာတ်ပုံ (သို့) ဗီဒီယိုထဲမှ အသုံးဝင်သော အချက်အလက်များကို ထုတ်ယူနိုင်အောင် လုပ်ဆောင်ပေးသည့် နယ်ပယ်တစ်ခု ဖြစ်ပါသည် — ဓာတ်ပုံထဲမှာ ဘာပါလဲ၊ object ဘယ်နေရာမှာရှိလဲ၊ video frame ထဲမှာ လှုပ်ရှားမှုရှိမရှိ စသည်တို့ကို ဆုံးဖြတ်ပေးခြင်းပင် ဖြစ်သည်။ ၎င်းသည် image processing (pixel data ကို ကိုင်တွယ်ခြင်း) နှင့် machine learning (နမူနာများမှ pattern များကို မှတ်သားခြင်း) ကြားထဲမှာ တည်ရှိပြီး၊ သမိုင်းအတော်များများမှာတော့ ဒီနှစ်ခုက သီးခြားစီပင် ဖြစ်ခဲ့ကြသည် — လူတစ်ဦးက pixel များကို feature ဟုခေါ်သော ကိန်းအနည်းငယ်အဖြစ် ပြောင်းလဲနည်းကို လက်ဖြင့် ဒီဇိုင်းဆွဲပေးပြီး၊ ရိုးရှင်းသော classifier တစ်ခုက ဆုံးဖြတ်ချက်ကို ချမှတ်ပေးခဲ့ကြသည်။ ဥပမာအားဖြင့် stop sign ဓာတ်ပုံတစ်ပုံကို 'အနီရောင်များပြီး ရှစ်ထောင့်ပုံသဏ္ဍာန်ရှိကာ S-T-O-P စာလုံးများပါဝင်သည်' ဟူသော လက်ဖြင့်ရေးသားထားသည့် စည်းမျဉ်းများဖြင့် လျှော့ချနိုင်သည် — ဤစည်းမျဉ်းတစ်ခုစီကို stop sign ကို stop sign ဖြစ်စေသည့် အချက်များအား လူတစ်ဦးက ဂရုတစိုက် စဉ်းစားပြီး ရွေးချယ်ခဲ့ခြင်း ဖြစ်သည်။

Deep Learning က ဤ pipeline ထဲမှာ လူက ဒီဇိုင်းဆွဲရသည့် အစိတ်အပိုင်းကို ပြောင်းလဲပစ်လိုက်သည်။ feature များကို လက်ဖြင့်ရွေးချယ်မည့်အစား၊ raw pixel များကို neural network တစ်ခုထံ တိုက်ရိုက်ပေးပို့ပြီး၊ label တပ်ထားသော နမူနာထောင်ပေါင်းများစွာမှ ဘယ် pixel pattern များကို အာရုံစိုက်သင့်သည်ကို ကိုယ်တိုင် သင်ယူခိုင်းလိုက်ခြင်းသာ ဖြစ်သည်။ Deep Learning with PyTorch သင်ခန်းစာမှာ network တစ်ခုက backpropagation ဖြင့် weight များကို သင်ယူပုံကို လေ့လာခဲ့ပြီးသားဖြစ်၍ ဤသင်ခန်းစာမှာတော့ အလားတူ mechanism ကို images များအပေါ် အထူးသက်ဆိုင်စေရန် ရည်ညွှန်းပြီး၊ color space, filtering, augmentation, spatial structure ကဲ့သို့ image-specific idea များ — images များကို သီးခြား toolkit လိုအပ်စေသည့် input အမျိုးအစားတစ်ခု ဖြစ်စေသည့် အကြောင်းရင်းများကို ဦးတည်ဆွေးနွေးမည် ဖြစ်သည်။

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

Tutorial Platform ပေါ်တွင် course တစ်ခုစီ၏ landing card မှာ banner image တစ်ပုံ လိုအပ်ပြီး၊ upload pipeline က ၎င်းကို visitor ထောင်နှင့်ချီအား ဝန်ဆောင်ပေးမီ မည်မျှပြင်းထန်စွာ compress လုပ်သင့်သည်ကို ဆုံးဖြတ်ရသည် — line-art diagram ရိုးရိုးလေးတစ်ခုကို classical rule များ (color အနည်းငယ်၊ edge ထက်ခြင်း၊ file size သေးငယ်ခြင်း) ဖြင့် သန့်ရှင်းစွာ compress လုပ်နိုင်ပြီး၊ ဓာတ်ပုံအမှန် banner တစ်ခုကတော့ 'ဒါဓာတ်ပုံပါ' ဆိုတာကို ယုံကြည်စိတ်ချရအောင် သိရှိရန် learned image classifier တစ်ခု လိုအပ်ကာ quality ထိန်းသိမ်းသည့် compression path ကွဲသို့ ပို့ဆောင်ပေးရသည် — ဤသင်ခန်းစာက မိတ်ဆက်ပေးသည့် classical-vs-learned ခွဲခြားမှုအား scale သေးငယ်သော ဥပမာကောင်းတစ်ခု ဖြစ်သည်။

အတူတူ စမ်းရေးကြည့်မယ်

text
CLASSICAL COMPUTER VISION PIPELINE
-----------------------------------
Raw Pixels  -->  Preprocess   -->  Hand-designed     -->  Hand-tuned      -->  Task Output
(H x W x 3)      (blur, gray,      Feature Extractor      Classifier           (label / count /
                  normalize)       (Sobel edges,          (SVM, decision        detection)
                                    SIFT keypoints,        tree, rules)
                                    HOG descriptors)

Every arrow above is designed by a human. A person decided which
edge kernel to use, which threshold counts as a "corner", and which
rule separates a cat from a dog.

DEEP-LEARNING COMPUTER VISION PIPELINE
---------------------------------------
Raw Pixels  ---------------------------------------->  Task Output
(H x W x 3)     Learned Feature Extractor                (label / count /
                (stacked Conv2d + pooling layers,         detection)
                 trained end-to-end by backprop)

Here a single trainable model replaces the hand-designed
middle stages. The network decides for itself which patterns
of pixels are worth detecting -- we only supply labeled
examples and a loss function.

THIS COURSE'S ROADMAP
----------------------
Basic         - this chapter. Images as tensors, filtering, transforms,
                classical features, and a CNN refresher aimed at vision.
Intermediate  - vision-specific architectures and training tricks:
                transfer learning, fine-tuning pretrained backbones,
                data pipelines for real image datasets.
Advanced      - object detection, semantic/instance segmentation,
                and other structured vision outputs beyond a single label.
Projects      - end-to-end builds combining earlier chapters into a
                deployable vision feature.
Exercises     - standalone practice problems across all difficulty levels.
You should see
ဒါသည် ASCII diagram သဘောတရားပုံဖြစ်ပြီး run နိုင်သော code မဟုတ်သဖြင့် program output ဟူ၍ မရှိပါ။ အပေါ်မှအောက်ဖတ်လျှင် — classical pipeline တွင် raw pixel နှင့် task output ကြားမှာ လူဒီဇိုင်းဆွဲထားသော အဆင့်လေးဆင့်ကို တွေ့ရမည်ဖြစ်ပြီး၊ deep-learning pipeline ကတော့ အလယ်ကအဆင့်များအားလုံးကို learned block တစ်ခုတည်းအဖြစ် ချုံ့ထားသည်၊ roadmap အပိုင်းကတော့ ဤသင်ခန်းစာကို Basic (ယခုအခန်း) မှ Exercises အထိ chapter ငါးခုဖြင့် စီစဉ်ထားကြောင်း ဖော်ပြထားသည်။

၅ မိနစ် စမ်းကြည့်

အထက်ပါ ASCII diagram ကို extend လုပ်ပြီး 'fixed input size ဆီ resize လုပ်ခြင်း' (သို့) 'pixel value များကို normalize လုပ်ခြင်း' ကဲ့သို့ အဆင့်တစ်ခုသည် pipeline တစ်ခုစီအတွင်း ဘယ်နေရာမှာရှိသင့်သလဲ ဆိုတာကို မှတ်စုရေးကြည့်ပါ — ၎င်းသည် preprocessing အပိုင်းလား၊ feature extraction အပိုင်းလား၊ (သို့) classical/learned ဘယ်ဟာမဆို ဖြစ်ပေါ်တတ်သော အဆင့်တစ်ခုလား။

သတိလေးတစ်ချက်

Classical CV technique များကို ခေတ်နောက်ကျပြီဟု ယူဆကာ လုံးဝကျော်သွားခြင်း — production pipeline အများစုသည် denoising, thresholding, color-space conversion ကဲ့သို့ classical အဆင့်များကို learned model ၏ ပတ်ဝန်းကျင်တွင် ဈေးသက်သာသော pre/post-processing အဖြစ် ယနေ့တိုင် အသုံးပြုနေဆဲဖြစ်၍ ၎င်းတို့ကို နားမလည်ပါက hybrid system များကို debug လုပ်ရန် ခက်ခဲစေသည်။

'Deep Learning က pipeline တစ်ခုလုံးကို အစားထိုးသည်' ဆိုသည်မှာ CNN တစ်ခုသည် preprocessing လုံးဝမလိုဟု အဓိပ္ပာယ်ဖွင့်ဆိုမိခြင်း — image pipeline အမှန်တွင် learned stage မမတိုင်မီ resize, normalize, augment လုပ်ငန်းစဉ်များ ဆက်လက်လိုအပ်ဆဲဖြစ်ပြီး ဤအဆင့်ကို ကျော်သွားခြင်းသည် beginner များ မကြာခဏ ကျူးလွန်တတ်သည့် အမှားဖြစ်ကာ accuracy ကို တိတ်တဆိတ် ထိခိုက်စေတတ်သည်။

Wikipedia — Computer visionComputer Vision

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • Classical CV technique များကို ခေတ်နောက်ကျပြီဟု ယူဆကာ လုံးဝကျော်သွားခြင်း — production pipeline အများစုသည် denoising, thresholding, color-space conversion ကဲ့သို့ classical အဆင့်များကို learned model ၏ ပတ်ဝန်းကျင်တွင် ဈေးသက်သာသော pre/post-processing အဖြစ် ယနေ့တိုင် အသုံးပြုနေဆဲဖြစ်၍ ၎င်းတို့ကို နားမလည်ပါက hybrid system များကို debug လုပ်ရန် ခက်ခဲစေသည်။
  • 'Deep Learning က pipeline တစ်ခုလုံးကို အစားထိုးသည်' ဆိုသည်မှာ CNN တစ်ခုသည် preprocessing လုံးဝမလိုဟု အဓိပ္ပာယ်ဖွင့်ဆိုမိခြင်း — image pipeline အမှန်တွင် learned stage မမတိုင်မီ resize, normalize, augment လုပ်ငန်းစဉ်များ ဆက်လက်လိုအပ်ဆဲဖြစ်ပြီး ဤအဆင့်ကို ကျော်သွားခြင်းသည် beginner များ မကြာခဏ ကျူးလွန်တတ်သည့် အမှားဖြစ်ကာ accuracy ကို တိတ်တဆိတ် ထိခိုက်စေတတ်သည်။
  • နမူနာ code ကို production system ပေါ် တိုက်ရိုက်မစမ်းဘဲ local/test environment တွင် အရင်အတည်ပြုပါ။

လေ့ကျင့်ခန်း

အထက်ပါ ASCII diagram ကို extend လုပ်ပြီး 'fixed input size ဆီ resize လုပ်ခြင်း' (သို့) 'pixel value များကို normalize လုပ်ခြင်း' ကဲ့သို့ အဆင့်တစ်ခုသည် pipeline တစ်ခုစီအတွင်း ဘယ်နေရာမှာရှိသင့်သလဲ ဆိုတာကို မှတ်စုရေးကြည့်ပါ — ၎င်းသည် preprocessing အပိုင်းလား၊ feature extraction အပိုင်းလား၊ (သို့) classical/learned ဘယ်ဟာမဆို ဖြစ်ပေါ်တတ်သော အဆင့်တစ်ခုလား။

You'll know it worked when: ဒါသည် ASCII diagram သဘောတရားပုံဖြစ်ပြီး run နိုင်သော code မဟုတ်သဖြင့် program output ဟူ၍ မရှိပါ။ အပေါ်မှအောက်ဖတ်လျှင် — classical pipeline တွင် raw pixel နှင့် task output ကြားမှာ လူဒီဇိုင်းဆွဲထားသော အဆင့်လေးဆင့်ကို တွေ့ရမည်ဖြစ်ပြီး၊ deep-learning pipeline ကတော့ အလယ်ကအဆင့်များအားလုံးကို learned block တစ်ခုတည်းအဖြစ် ချုံ့ထားသည်၊ roadmap အပိုင်းကတော့ ဤသင်ခန်းစာကို Basic (ယခုအခန်း) မှ Exercises အထိ chapter ငါးခုဖြင့် စီစဉ်ထားကြောင်း ဖော်ပြထားသည်။

Computer Vision ဆိုတာ ဘာလဲ | Thuta Learning