Thuta Learning
Local AI / Local LLM
ExercisesAIbeginner

လေ့ကျင့်ခန်း: နှေးကွေးနေသော Local Setup တစ်ခုကို Diagnose ပြုလုပ်ခြင်း

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • လေ့ကျင့်ခန်း: နှေးကွေးနေသော Local Setup တစ်ခုကို Diagnose ပြုလုပ်ခြင်း concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • Diagram ကို ဖတ်ပြီး architecture ထဲမှာ data/request ဘယ်လိုစီးဆင်းသလဲ ခြေရာခံနိုင်ရန်
  • ကိုယ့် hardware/use case အတွက် ဘယ်လို ရွေးချယ်သင့်သလဲ ဆုံးဖြတ်နိုင်ရန်

နားလည်ထားရမယ့် အချက်

Local AI setup တစ်ခု နှေးကွေးနေခြင်း (သို့) အသုံးမပြုနိုင်လောက်အောင် ဖြစ်နေခြင်းကို diagnose လုပ်တတ်ခြင်းသည် runtime install လုပ်တတ်ခြင်း (သို့) model download လုပ်တတ်ခြင်းနှင့် လုံးဝမတူသည့် ကျွမ်းကျင်မှုတစ်ခုဖြစ်သည်။ အဓိက diagnostic မေးခွန်းသည် အမြဲတမ်း matching ပြဿနာတစ်ခုသာဖြစ်သည် - သင့်တွင်ရှိသော hardware သည် သင်ရွေးချယ်ထားသော model size နှင့် quantization level ကိုကိုက်ညီပါသလား၊ model ၏ capability သည် သင်တောင်းဆိုနေသော task ၏ ခက်ခဲမှုအဆင့်ကို ကိုက်ညီပါသလား။

လက္ခဏာ (symptom) နှစ်မျိုးက mismatch နှစ်မျိုးကို ညွှန်ပြသည်။ Model တစ်ခု load ရန်နှေးခြင်း၊ crash ဖြစ်ခြင်း (သို့) disk ကို swap လုပ်နေရင်း machine တစ်ခုလုံး ရပ်တန့်သွားခြင်းများသည် memory mismatch ဖြစ်ဖွယ်များသည် - သင် download လုပ်ထားသော precision အတိုင်း model weight များသည် ရရှိနိုင်သော RAM (သို့) VRAM ထဲသို့ လုံးဝမကိုက်ညီခြင်းဖြစ်သည်။

ဒါမှမဟုတ် model က load ဖြစ်ပြီး မြန်မြန်ဆန်ဆန် တုံ့ပြန်သော်လည်း အဖြေများသည် ပေါ့ပေါ့ပါးပါး၊ မှားယွင်း (သို့) ထပ်ခါထပ်ခါဖြစ်နေပါက capability mismatch ဖြစ်သည် - task ၏ reasoning လိုအပ်ချက်နှင့်နှိုင်းလျှင် model သည် သေးလွန်း (သို့) quantize အလွန်အကျွံလုပ်ထားသည်ကို ရွေးချယ်မိခြင်းဖြစ်သည်။

ဒီနှစ်မျိုးအတွက် ပြင်ဆင်ချက်ကွာခြားသောကြောင့် setting များကို မထိမီ ဘယ်တစ်ခု ကြုံနေရသည်ကို ဦးစွာသိထားရန် အရေးကြီးသည်။ Basic အခန်းမှ estimation method (parameter count ကို bytes-per-parameter နှင့်မြှောက်ပြီး gigabyte သို့ ပြောင်းခြင်း) သည် memory mismatch အမျိုးအစားကို ဖြေရှင်းရန် tool တစ်ခုဖြစ်သည် - "နှေးနေတယ်" ဆိုသော မရေရာသော လက္ခဏာကို သင့် hardware နှင့်တိုက်ရိုက်နှိုင်းယှဉ်နိုင်သည့် တိကျသော ဂဏန်းအဖြစ် ပြောင်းလဲပေးသည်။

ဒီ lesson က memory mismatch ဖြစ်စဉ်တစ်ခုကို အသေးစိတ်လေ့လာစေပြီး၊ ဖတ်ရုံမျှမက ကိုယ်တိုင်တွက်ချက်ကြည့်စေမည်ဖြစ်သည်။ Run မလုပ်ခင် ခန့်မှန်းဆိုသော အလေ့အကျင့်သည် သင့် local AI stack ထဲတွင် မမျှော်လင့်သလို တစ်ခုခု ဖြစ်ပေါ်လာသည့်အခါ ယုံကြည်စိတ်ချစွာ ပြဿနာဖြေရှင်းခြင်းနှင့် စမ်းသပ်ခန်းမှန်း guessing တို့ကို ခွဲခြားပေးသည့်အချက်ပင်ဖြစ်သည်။

text
MEMORY MISMATCH: 70B MODEL VS 16 GB LAPTOP RAM
----------------------------------------------
MEMORY MISMATCH: 70B MODEL AT FP16 VS 16 GB LAPTOP RAM
-----------------------------------------------------------
REQUIRED  (70B params x 2 bytes/param, converted to GB)
  [########################################] 130.4 GB

AVAILABLE (laptop RAM, no dedicated GPU)
  [#####]                                     16.0 GB

  0                                              140 GB
  |----|----|----|----|----|----|----|----|----|

GAP: about 114.4 GB more than the laptop actually has

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

လက်တွေ့တွင် ဒီ diagnostic အလေ့အကျင့်က ဘုံ trap နှစ်ခုကို ရှောင်ရှားပေးနိုင်သည်။ ပထမ trap က trial-and-error ဖြင့် hardware ဝယ်ယူခြင်းဖြစ်သည် - memory ဟာ bottleneck အမှန်ဖြစ်ကြောင်း အတည်မပြုမီ RAM ပိုထည့်ခြင်း (သို့) GPU ကြီးတစ်ခုဝယ်ခြင်း။

ငွေမကုန်ခင် ငါးမိနစ်လောက်ကြာသော တွက်ချက်မှုတစ်ခုက quantization level လျှော့ချခြင်း (သို့) model သေးလိုက်ခြင်းကဲ့သို့ ပိုသက်သာသော ဖြေရှင်းချက်တစ်ခုက ဒီပန်းတိုင်ကိုရောက်စေမလားဆိုတာကို တိကျစွာပြောပြနိုင်သည်။

ဒုတိယ trap ကတော့ stack ၏ layer မှားကို အပြစ်တင်ခြင်းဖြစ်သည် - memory ထဲတွင် ကြပ်တင်းစွာသာ ကိုက်ညီသော model တစ်ခုသည် install "bug ရှိသည်" (သို့) runtime "နှေးသည်" ဟု ထင်ရနိုင်သော်လည်း၊ RAM ထက် များစွာနှေးသော disk ဆီသို့ operating system က model weight များကို အမြဲတမ်း swap လုပ်နေခြင်းသာဖြစ်သည်။ Estimate ကို ဦးစွာ run ကြည့်ခြင်းက configuration file များကို စတင်ချိန်ညှိမီ တရားမျှတသော ပြိုင်ပွဲတစ်ခုထဲ ရှိမရှိကို ပြောပြပေးသည်။

Parameter Count ရှာပါ

သင်စဉ်းစားနေသော model ၏ parameter count ကို ရှာပါ။

Quantization Level ဆုံးဖြတ်ပါ

မည်သည့် quantization level ကို အသုံးပြုမည်ကို ဆုံးဖြတ်ပါ။

တွက်ချက်မှု ပြန်လုပ်ပါ

ဒီ lesson ထဲက တွက်ချက်မှုအတူတူကို ထပ်လုပ်ပါ။

သင့် Hardware အစစ်နှင့် နှိုင်းယှဉ်ပါ

ရလဒ်ကို သင်တကယ်ရှိသော RAM (သို့) VRAM နှင့် နှိုင်းယှဉ်ပါ၊ သင်ဆန္ဒရှိသော RAM မဟုတ်ပါ။

ရလဒ်ဂဏန်းကို အနိမ့်ဆုံးအခြေခံအဖြစ်သာ သဘောထားပါ၊ အာမခံချက်မဟုတ်ပါ - operating system၊ အခြား run နေသော application များနှင့် runtime ကိုယ်တိုင်၏ overhead တို့ကလည်း memory pool တူတူထဲမှ စားသုံးကြသောကြောင့်ဖြစ်သည်။

အတူတူ စမ်းရေးကြည့်မယ်

python
def estimate_memory_gb(param_count_billions, bytes_per_param):
    """Estimate RAM/VRAM needed to load a model's weights.

    param_count_billions: number of parameters, in billions (e.g. 70 for
                           a 70B-parameter model)
    bytes_per_param: bytes used to store each parameter (2 for FP16/BF16
                      full precision, 1 for 8-bit quantization, ~0.5 for
                      4-bit quantization)
    Returns the estimated size in gigabytes (GiB).
    """
    params = param_count_billions * 1_000_000_000
    total_bytes = params * bytes_per_param
    return total_bytes / (1024 ** 3)


def check_fit(required_gb, available_gb):
    fits = required_gb <= available_gb
    shortfall = max(0.0, required_gb - available_gb)
    return fits, shortfall


# The scenario: a 70B-parameter model downloaded at FP16 (2 bytes/param),
# on a laptop with 16 GB of RAM and no dedicated GPU.
required = estimate_memory_gb(param_count_billions=70, bytes_per_param=2.0)
available = 16
fits, shortfall = check_fit(required, available)

print(f"Estimated requirement: {required:.1f} GB")
print(f"Available RAM:         {available} GB")
print(f"Fits in available memory? {fits}")
print(f"Shortfall: {shortfall:.1f} GB")
You should see
Function ကို scenario ပေါ်တွင် run လိုက်လျှင် ဤအတိုင်း print ထုတ်သည် - estimated requirement 130.4 GB (parameter ဘီလီယံ 70 x bytes 2 ကို gigabyte ပြောင်းထားသည်)၊ available RAM 16 GB၊ `Fits in available memory? False`၊ နှင့် shortfall 114.4 GB -- ၎င်းက ဒါဟာ အနည်းငယ် နှေးခြင်းမျှမက ပြင်းထန်သော memory mismatch တစ်ခုဖြစ်ကြောင်း အတည်ပြုပေးသည်။

၅ မိနစ် စမ်းကြည့်

လုပ်ဖော်ကိုင်ဖက်တစ်ဦးက 70B-parameter open-weights model တစ်ခုကို FP16 (parameter တစ်ခုလျှင် bytes 2) ဖြင့် RAM 16 GB ရှိပြီး dedicated GPU မရှိသော laptop ပေါ်သို့ download လုပ်လိုက်သည်။ Load ဖြစ်ရန် မိနစ်များစွာကြာပြီး၊ fan များအပြည့်အဝလည်နေကာ၊ နောက်ဆုံးအဖြေပြန်ရသောအခါ token တစ်ခုစီအတွက် စက္ကန့်အနည်းငယ်ကြာသည်။ တစ်စုံတစ်ခု မပြောင်းလဲမီ၊ ဒီ lesson ထဲက method ကိုသုံးပြီး FP16 ဖြင့် ဒီ model ၏ estimated memory requirement ကို တွက်ချက်ပြီး available 16 GB နှင့် နှိုင်းယှဉ်ပါ။ ထို့နောက် 16 GB RAM ထဲသို့ တကယ်ကိုက်ညီမည့် alternative configuration နှစ်ခုကို ရှာပါ - တစ်ခုက 70B model အတိုင်းထားပြီး quantization level ပြောင်းခြင်း၊ နောက်တစ်ခုက FP16 precision ကို ဆက်ထားပြီး parameter count ပြောင်းခြင်း။ Alternative တစ်ခုချင်းစီအတွက် ရလဒ် estimated memory requirement ကို ဖော်ပြပါ။ ဒီ alternative နှစ်ခုထဲက ဘယ်တစ်ခုကို အကြံပြုမလဲ၊ ဘာကြောင့်လဲ -- ကိုက်ညီမှုသာမက quality ဘယ်လိုလျှော့ချရသည်ကိုပါ စဉ်းစားပါ။

သတိလေးတစ်ချက်

Estimate ဟာ အပြည့်အစုံဖြစ်သည်ဟု ယူဆမိခြင်း -- အမှန်တကယ် memory အသုံးပြုမှုတွင် KV cache၊ context length နှင့် runtime overhead ပါဝင်သောကြောင့် real usage သည် weight raw calculation ထက် များသောအားဖြင့် ပိုမြင့်တတ်သည်။

Quantization level ပြောင်းခြင်း (သို့) model သေးသေးလေးရွေးချယ်ခြင်းက ပြဿနာကို အခမဲ့ဖြေရှင်းနိုင်မလားဆိုတာကို ဦးစွာမစစ်ဆေးဘဲ hardware ပိုဝယ်ဖို့ ရုတ်တရက်ကူးလွန်မိခြင်း။

Hugging Face -- Model Memory CalculatorLocal AI / Local LLM

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • Estimate ဟာ အပြည့်အစုံဖြစ်သည်ဟု ယူဆမိခြင်း -- အမှန်တကယ် memory အသုံးပြုမှုတွင် KV cache၊ context length နှင့် runtime overhead ပါဝင်သောကြောင့် real usage သည် weight raw calculation ထက် များသောအားဖြင့် ပိုမြင့်တတ်သည်။
  • Quantization level ပြောင်းခြင်း (သို့) model သေးသေးလေးရွေးချယ်ခြင်းက ပြဿနာကို အခမဲ့ဖြေရှင်းနိုင်မလားဆိုတာကို ဦးစွာမစစ်ဆေးဘဲ hardware ပိုဝယ်ဖို့ ရုတ်တရက်ကူးလွန်မိခြင်း။
  • Model (သို့) tool အသစ်တစ်ခုကို production/daily-use workflow ထဲ တိုက်ရိုက်မထည့်ခင် သေးငယ်တဲ့ scale နဲ့ အရင်စမ်းကြည့်ပါ။

လေ့ကျင့်ခန်း

လုပ်ဖော်ကိုင်ဖက်တစ်ဦးက 70B-parameter open-weights model တစ်ခုကို FP16 (parameter တစ်ခုလျှင် bytes 2) ဖြင့် RAM 16 GB ရှိပြီး dedicated GPU မရှိသော laptop ပေါ်သို့ download လုပ်လိုက်သည်။ Load ဖြစ်ရန် မိနစ်များစွာကြာပြီး၊ fan များအပြည့်အဝလည်နေကာ၊ နောက်ဆုံးအဖြေပြန်ရသောအခါ token တစ်ခုစီအတွက် စက္ကန့်အနည်းငယ်ကြာသည်။ တစ်စုံတစ်ခု မပြောင်းလဲမီ၊ ဒီ lesson ထဲက method ကိုသုံးပြီး FP16 ဖြင့် ဒီ model ၏ estimated memory requirement ကို တွက်ချက်ပြီး available 16 GB နှင့် နှိုင်းယှဉ်ပါ။ ထို့နောက် 16 GB RAM ထဲသို့ တကယ်ကိုက်ညီမည့် alternative configuration နှစ်ခုကို ရှာပါ - တစ်ခုက 70B model အတိုင်းထားပြီး quantization level ပြောင်းခြင်း၊ နောက်တစ်ခုက FP16 precision ကို ဆက်ထားပြီး parameter count ပြောင်းခြင်း။ Alternative တစ်ခုချင်းစီအတွက် ရလဒ် estimated memory requirement ကို ဖော်ပြပါ။ ဒီ alternative နှစ်ခုထဲက ဘယ်တစ်ခုကို အကြံပြုမလဲ၊ ဘာကြောင့်လဲ -- ကိုက်ညီမှုသာမက quality ဘယ်လိုလျှော့ချရသည်ကိုပါ စဉ်းစားပါ။

You'll know it worked when: Function ကို scenario ပေါ်တွင် run လိုက်လျှင် ဤအတိုင်း print ထုတ်သည် - estimated requirement 130.4 GB (parameter ဘီလီယံ 70 x bytes 2 ကို gigabyte ပြောင်းထားသည်)၊ available RAM 16 GB၊ `Fits in available memory? False`၊ နှင့် shortfall 114.4 GB -- ၎င်းက ဒါဟာ အနည်းငယ် နှေးခြင်းမျှမက ပြင်းထန်သော memory mismatch တစ်ခုဖြစ်ကြောင်း အတည်ပြုပေးသည်။

လေ့ကျင့်ခန်း: နှေးကွေးနေသော Local Setup တစ်ခုကို Diagnose ပြုလုပ်ခြင်း | Thuta Learning