Thuta Learning
Local AI / Local LLM
BasicAIbeginner

Model Format များနှင့် Model ရွေးချယ်ခြင်း

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • Model Format များနှင့် Model ရွေးချယ်ခြင်း concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • Diagram ကို ဖတ်ပြီး architecture ထဲမှာ data/request ဘယ်လိုစီးဆင်းသလဲ ခြေရာခံနိုင်ရန်
  • ကိုယ့် hardware/use case အတွက် ဘယ်လို ရွေးချယ်သင့်သလဲ ဆုံးဖြတ်နိုင်ရန်

နားလည်ထားရမယ့် အချက်

model ရဲ့ architecture ဆိုတာ ဒီဇိုင်းပါ - layer တွေ ဘယ်လိုချိတ်ဆက်ထား လဲဆိုတာပါ - ဒါပေမယ့် ဒီဒီဇိုင်းကို disk ပေါ်မှာ ပုံစံတစ်ခုခုနဲ့ သိမ်းရဆဲပါ၊ ပုံစံတစ်မျိုးထက်ပိုပြီး ရှိပါတယ်။

ဒါက model file format ဆိုတာပါပဲ - model ရဲ့ weight နဲ့ structure ကို file တစ်ခုအဖြစ် package လုပ်ပုံအတိအကျဖြစ်ပြီး architecture ကိုယ်တိုင်နဲ့ သီးခြားပါ၊ ဓာတ်ပုံတစ်ခုကို JPEG (သို့) PNG အဖြစ် သိမ်းလို့ရသလို ပုံထဲက အကြောင်း အရာ မပြောင်းသလိုမျိုးပါ။

format သုံးမျိုး ရေပန်းစားပါတယ် -

  • GGUF - quantize လုပ်ထားတဲ့ model တွေကို consumer hardware ပေါ်မှာ efficient run ဖို့ တည်ဆောက်ထားပြီး beginner-friendly local tool အများစုက မျှော်လင့်ထားတဲ့ format
  • Safetensors - safety နဲ့ load speed ကို အဓိကထား၊ ဖွင့်တဲ့အခါ arbitrary code run လုပ်လို့ လုံးဝမရအောင် ဒီဇိုင်းလုပ်ထား
  • ONNX - portability အတွက် ဒီဇိုင်းလုပ်ထား၊ framework တစ်ခုမှာ train ခဲ့တဲ့ model ကို runtime မျိုးစုံမှာ၊ GPU မဟုတ်တဲ့ hardware ပေါ်မှာတောင် run စေနိုင်

model ဘယ်ဟာမှ မ download လုပ်ခင် အတွေ့အကြုံရှိတဲ့ user တွေက model card ကို ဖတ်ကြပါတယ် - model ဟာ ဘာလဲ၊ ဘယ်သူကလုပ်ခဲ့လဲ၊ license ဘယ်လိုသုံးပြီး regulate လုပ်ထားလဲ၊ limitation ဘာတွေရှိလဲဆိုတာ ဖော်ပြထားတဲ့ documentation page ဖြစ်ပြီး model file နဲ့ တွဲပြီး host လုပ်ထားလေ့ရှိပါတယ်။ model card ကို packaged food ပေါ်က label ကို ဆက်ဆံသလိုမျိုး ဆက်ဆံပါ - resource တွေ ပေးဆပ်မလုပ်ခင် အထဲမှာ တကယ်ဘာပါလဲဆိုတာ ပြောပြပါတယ်။

model card ကျော်ဖတ်ခြင်းက ခင်ဗျားမှာ မရှိတဲ့ hardware လိုအပ်တယ်ဆိုတာ (သို့) ခင်ဗျားလုပ်ချင်တဲ့ အသုံးပြုမှုကို တားမြစ်ထားတဲ့ license ပါမှန်းသိတဲ့အထိ gigabyte အများကြီး download ချပြီးမှ တွေ့ရနိုင်ပါတယ်။

GGUF
quantize လုပ်ထားတဲ့ model တွေကို consumer hardware ပေါ်မှာ efficient run ဖို့ တည်ဆောက်ထားတဲ့ file format ပါ။
Safetensors
safety နဲ့ load speed ကို အဓိကထားပြီး ဖွင့်တဲ့အခါ arbitrary code run လုပ်လို့မရအောင် ဒီဇိုင်းလုပ်ထားတဲ့ file format ပါ။
ONNX
framework တစ်ခုမှာ train ထားတဲ့ model ကို runtime မျိုးစုံမှာ run နိုင်စေတဲ့ portable format ပါ။
Model Card
model ရဲ့ ဖန်တီးသူ၊ license၊ limitation တွေကို ဖော်ပြထားတဲ့ documentation page ပါ။
text
FROM REPOSITORY TO RUNNING MODEL
--------------------------------
[Model Repository]
        |
  [Model Card]  <- read this FIRST: license, size, format
        |
  [Download File]  (.gguf / .safetensors / .onnx)
        |
  [Local Runtime]  loads the file and runs inference

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

ခင်ဗျားရဲ့ ပထမဆုံး local model ကို download လုပ်တော့မယ်ဆိုပြီး name တူတူ extension မတူတဲ့ file သုံးခု တွေ့ရတယ်ဆိုပါစို့ - .gguf တစ်ခု၊ .safetensors တစ်ခု။ မှန်းဆမနေဘဲ ခင်ဗျားရွေးထားတဲ့ local tool က ဘာကို တကယ်မျှော်လင့်လဲ စစ်ပါ - GGUF-based runtime ရိုးရိုး beginner tool အများစုက .gguf file ကို တိတိကျကျ လိုချင်ပြီး Python-based workflow တွေကတော့ Safetensors ကို မျှော်လင့်လေ့ ရှိပါတယ်။

ဘယ်ဟာမှ မ download လုပ်ခင် repository page ပေါ်က model card ကို ဖွင့်ပြီး - license ဘာလဲ၊ ခင်ဗျား VRAM နဲ့ ကိုက်ညီလား၊ ခင်ဗျား လုပ်ချင်တဲ့ task အတွက် တကယ်ရည်ရွယ်ထားလားဆိုတာ စစ်ပါ။

ဒီ ငါးမိနစ်လောက် စစ်ဆေးမှုက နာရီအများကြီး ချွေတာပေးပါတယ် - bandwidth နဲ့ disk space သုံးပြီးမှ commercial use ကို တားမြစ်ထားတဲ့ license ပါတဲ့ model (business အတွက် လိုချင်နေတဲ့အချိန်) (သို့) VRAM 24GB လိုအပ်တဲ့ model (8GB ပဲရှိတဲ့အချိန်) ကို မတွေ့မီ ကြိုတင်တားဆီးပေး ပါတယ်။ အောက်က code ဟာ ဒီစစ်ဆေးမှုရဲ့ အစိတ်အပိုင်းငယ်တစ်ခုကို automate လုပ်ပါတယ် - model card မှာ ခင်ဗျားတကယ်လိုအပ်တဲ့ field တွေ ရှိမရှိ ယုံကြည်ခင် အတည်ပြုတာပါ။

Model Download လုပ်ခင် စစ်ရန်

အတူတူ စမ်းရေးကြည့်မယ်

python
REQUIRED_FIELDS = ["name", "params", "format", "license", "quantization"]


def check_model_card(card):
    """Check a model card dict for the fields a beginner should
    always verify before downloading a model."""
    present = [f for f in REQUIRED_FIELDS if f in card]
    missing = [f for f in REQUIRED_FIELDS if f not in card]
    return present, missing


cards = {
    "complete-example": {
        "name": "Llama-3-8B-Instruct",
        "params": "8B",
        "format": "GGUF",
        "license": "Llama 3 Community License",
        "quantization": "Q4_K_M",
    },
    "incomplete-example": {
        "name": "MysteryModel-7B",
        "params": "7B",
        "format": "safetensors",
    },
}

for card_name, card in cards.items():
    present, missing = check_model_card(card)
    print(f"{card_name}:")
    print(f"  present: {present}")
    print(f"  missing: {missing}")
You should see
model card dict နှစ်ခုကို required field list နဲ့ စစ်ဆေးပြီး present/missing field တွေ ပြထားပါတယ်။ Output အတိအကျမှာ - complete-example:
  present: ['name', 'params', 'format', 'license', 'quantization']
  missing: []
incomplete-example:
  present: ['name', 'params', 'format']
  missing: ['license', 'quantization']

၅ မိနစ် စမ်းကြည့်

REQUIRED_FIELDS list ထဲကို "context_window" အသစ်တစ်ခု ထည့်ပြီး ပြန် run ကြည့်ပါ။ cards dict ထဲကို ခင်ဗျားကိုယ်ပိုင် model card တစ်ခု ထည့်ပြီး missing field ဘာတွေရှိလဲ ကြည့်ပါ။

သတိလေးတစ်ချက်

file format (GGUF/Safetensors) ကို model architecture နဲ့ ရောထွေးမိပြီး format တူရင် model တူမယ်လို့ ထင်ခြင်း

download မလုပ်ခင် model card ကို မဖတ်ဘဲ ကျော်ပြီး license (သို့) hardware requirement ကို နောက်မှသာ တွေ့ခြင်း

Hugging Face docs: GGUFLocal AI / Local LLM

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • file format (GGUF/Safetensors) ကို model architecture နဲ့ ရောထွေးမိပြီး format တူရင် model တူမယ်လို့ ထင်ခြင်း
  • download မလုပ်ခင် model card ကို မဖတ်ဘဲ ကျော်ပြီး license (သို့) hardware requirement ကို နောက်မှသာ တွေ့ခြင်း
  • Model (သို့) tool အသစ်တစ်ခုကို production/daily-use workflow ထဲ တိုက်ရိုက်မထည့်ခင် သေးငယ်တဲ့ scale နဲ့ အရင်စမ်းကြည့်ပါ။

လေ့ကျင့်ခန်း

REQUIRED_FIELDS list ထဲကို "context_window" အသစ်တစ်ခု ထည့်ပြီး ပြန် run ကြည့်ပါ။ cards dict ထဲကို ခင်ဗျားကိုယ်ပိုင် model card တစ်ခု ထည့်ပြီး missing field ဘာတွေရှိလဲ ကြည့်ပါ။

You'll know it worked when: model card dict နှစ်ခုကို required field list နဲ့ စစ်ဆေးပြီး present/missing field တွေ ပြထားပါတယ်။ Output အတိအကျမှာ - complete-example: present: ['name', 'params', 'format', 'license', 'quantization'] missing: [] incomplete-example: present: ['name', 'params', 'format'] missing: ['license', 'quantization']

Model Format များနှင့် Model ရွေးချယ်ခြင်း | Thuta Learning