နားလည်ထားရမယ့် အချက်
model ရဲ့ architecture ဆိုတာ ဒီဇိုင်းပါ - layer တွေ ဘယ်လိုချိတ်ဆက်ထား လဲဆိုတာပါ - ဒါပေမယ့် ဒီဒီဇိုင်းကို disk ပေါ်မှာ ပုံစံတစ်ခုခုနဲ့ သိမ်းရဆဲပါ၊ ပုံစံတစ်မျိုးထက်ပိုပြီး ရှိပါတယ်။
ဒါက model file format ဆိုတာပါပဲ - model ရဲ့ weight နဲ့ structure ကို file တစ်ခုအဖြစ် package လုပ်ပုံအတိအကျဖြစ်ပြီး architecture ကိုယ်တိုင်နဲ့ သီးခြားပါ၊ ဓာတ်ပုံတစ်ခုကို JPEG (သို့) PNG အဖြစ် သိမ်းလို့ရသလို ပုံထဲက အကြောင်း အရာ မပြောင်းသလိုမျိုးပါ။
format သုံးမျိုး ရေပန်းစားပါတယ် -
- GGUF - quantize လုပ်ထားတဲ့ model တွေကို consumer hardware ပေါ်မှာ efficient run ဖို့ တည်ဆောက်ထားပြီး beginner-friendly local tool အများစုက မျှော်လင့်ထားတဲ့ format
- Safetensors - safety နဲ့ load speed ကို အဓိကထား၊ ဖွင့်တဲ့အခါ arbitrary code run လုပ်လို့ လုံးဝမရအောင် ဒီဇိုင်းလုပ်ထား
- ONNX - portability အတွက် ဒီဇိုင်းလုပ်ထား၊ framework တစ်ခုမှာ train ခဲ့တဲ့ model ကို runtime မျိုးစုံမှာ၊ GPU မဟုတ်တဲ့ hardware ပေါ်မှာတောင် run စေနိုင်
model ဘယ်ဟာမှ မ download လုပ်ခင် အတွေ့အကြုံရှိတဲ့ user တွေက model card ကို ဖတ်ကြပါတယ် - model ဟာ ဘာလဲ၊ ဘယ်သူကလုပ်ခဲ့လဲ၊ license ဘယ်လိုသုံးပြီး regulate လုပ်ထားလဲ၊ limitation ဘာတွေရှိလဲဆိုတာ ဖော်ပြထားတဲ့ documentation page ဖြစ်ပြီး model file နဲ့ တွဲပြီး host လုပ်ထားလေ့ရှိပါတယ်။ model card ကို packaged food ပေါ်က label ကို ဆက်ဆံသလိုမျိုး ဆက်ဆံပါ - resource တွေ ပေးဆပ်မလုပ်ခင် အထဲမှာ တကယ်ဘာပါလဲဆိုတာ ပြောပြပါတယ်။
model card ကျော်ဖတ်ခြင်းက ခင်ဗျားမှာ မရှိတဲ့ hardware လိုအပ်တယ်ဆိုတာ (သို့) ခင်ဗျားလုပ်ချင်တဲ့ အသုံးပြုမှုကို တားမြစ်ထားတဲ့ license ပါမှန်းသိတဲ့အထိ gigabyte အများကြီး download ချပြီးမှ တွေ့ရနိုင်ပါတယ်။
- GGUF
- quantize လုပ်ထားတဲ့ model တွေကို consumer hardware ပေါ်မှာ efficient run ဖို့ တည်ဆောက်ထားတဲ့ file format ပါ။
- Safetensors
- safety နဲ့ load speed ကို အဓိကထားပြီး ဖွင့်တဲ့အခါ arbitrary code run လုပ်လို့မရအောင် ဒီဇိုင်းလုပ်ထားတဲ့ file format ပါ။
- ONNX
- framework တစ်ခုမှာ train ထားတဲ့ model ကို runtime မျိုးစုံမှာ run နိုင်စေတဲ့ portable format ပါ။
- Model Card
- model ရဲ့ ဖန်တီးသူ၊ license၊ limitation တွေကို ဖော်ပြထားတဲ့ documentation page ပါ။
FROM REPOSITORY TO RUNNING MODEL
--------------------------------
[Model Repository]
|
[Model Card] <- read this FIRST: license, size, format
|
[Download File] (.gguf / .safetensors / .onnx)
|
[Local Runtime] loads the file and runs inferenceလက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်
ခင်ဗျားရဲ့ ပထမဆုံး local model ကို download လုပ်တော့မယ်ဆိုပြီး name တူတူ extension မတူတဲ့ file သုံးခု တွေ့ရတယ်ဆိုပါစို့ - .gguf တစ်ခု၊ .safetensors တစ်ခု။ မှန်းဆမနေဘဲ ခင်ဗျားရွေးထားတဲ့ local tool က ဘာကို တကယ်မျှော်လင့်လဲ စစ်ပါ - GGUF-based runtime ရိုးရိုး beginner tool အများစုက .gguf file ကို တိတိကျကျ လိုချင်ပြီး Python-based workflow တွေကတော့ Safetensors ကို မျှော်လင့်လေ့ ရှိပါတယ်။
ဘယ်ဟာမှ မ download လုပ်ခင် repository page ပေါ်က model card ကို ဖွင့်ပြီး - license ဘာလဲ၊ ခင်ဗျား VRAM နဲ့ ကိုက်ညီလား၊ ခင်ဗျား လုပ်ချင်တဲ့ task အတွက် တကယ်ရည်ရွယ်ထားလားဆိုတာ စစ်ပါ။
ဒီ ငါးမိနစ်လောက် စစ်ဆေးမှုက နာရီအများကြီး ချွေတာပေးပါတယ် - bandwidth နဲ့ disk space သုံးပြီးမှ commercial use ကို တားမြစ်ထားတဲ့ license ပါတဲ့ model (business အတွက် လိုချင်နေတဲ့အချိန်) (သို့) VRAM 24GB လိုအပ်တဲ့ model (8GB ပဲရှိတဲ့အချိန်) ကို မတွေ့မီ ကြိုတင်တားဆီးပေး ပါတယ်။ အောက်က code ဟာ ဒီစစ်ဆေးမှုရဲ့ အစိတ်အပိုင်းငယ်တစ်ခုကို automate လုပ်ပါတယ် - model card မှာ ခင်ဗျားတကယ်လိုအပ်တဲ့ field တွေ ရှိမရှိ ယုံကြည်ခင် အတည်ပြုတာပါ။
Model Download လုပ်ခင် စစ်ရန်
အတူတူ စမ်းရေးကြည့်မယ်
REQUIRED_FIELDS = ["name", "params", "format", "license", "quantization"]
def check_model_card(card):
"""Check a model card dict for the fields a beginner should
always verify before downloading a model."""
present = [f for f in REQUIRED_FIELDS if f in card]
missing = [f for f in REQUIRED_FIELDS if f not in card]
return present, missing
cards = {
"complete-example": {
"name": "Llama-3-8B-Instruct",
"params": "8B",
"format": "GGUF",
"license": "Llama 3 Community License",
"quantization": "Q4_K_M",
},
"incomplete-example": {
"name": "MysteryModel-7B",
"params": "7B",
"format": "safetensors",
},
}
for card_name, card in cards.items():
present, missing = check_model_card(card)
print(f"{card_name}:")
print(f" present: {present}")
print(f" missing: {missing}")
model card dict နှစ်ခုကို required field list နဲ့ စစ်ဆေးပြီး present/missing field တွေ ပြထားပါတယ်။ Output အတိအကျမှာ - complete-example:
present: ['name', 'params', 'format', 'license', 'quantization']
missing: []
incomplete-example:
present: ['name', 'params', 'format']
missing: ['license', 'quantization']၅ မိနစ် စမ်းကြည့်
REQUIRED_FIELDS list ထဲကို "context_window" အသစ်တစ်ခု ထည့်ပြီး ပြန် run ကြည့်ပါ။ cards dict ထဲကို ခင်ဗျားကိုယ်ပိုင် model card တစ်ခု ထည့်ပြီး missing field ဘာတွေရှိလဲ ကြည့်ပါ။
သတိလေးတစ်ချက်
file format (GGUF/Safetensors) ကို model architecture နဲ့ ရောထွေးမိပြီး format တူရင် model တူမယ်လို့ ထင်ခြင်း
download မလုပ်ခင် model card ကို မဖတ်ဘဲ ကျော်ပြီး license (သို့) hardware requirement ကို နောက်မှသာ တွေ့ခြင်း
Hugging Face docs: GGUF — Local AI / Local LLM