Thuta Learning
Local AI / Local LLM
BasicAIbeginner

Hardware အခြေခံများ - CPU, GPU, RAM, VRAM

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • Hardware အခြေခံများ - CPU, GPU, RAM, VRAM concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • Diagram ကို ဖတ်ပြီး architecture ထဲမှာ data/request ဘယ်လိုစီးဆင်းသလဲ ခြေရာခံနိုင်ရန်
  • ကိုယ့် hardware/use case အတွက် ဘယ်လို ရွေးချယ်သင့်သလဲ ဆုံးဖြတ်နိုင်ရန်

နားလည်ထားရမယ့် အချက်

model file တစ်ခုက hard drive ပေါ်မှာ ထိုင်နေရုံနဲ့ အသုံးမဝင်ပါဘူး - တစ်ခုခုက ၎င်းကို working memory ထဲကို load လုပ်ပြီး တွက်ချက်မှု အမှန်တကယ် လုပ်ပေးရပါမယ်။ ဒီ 'တစ်ခုခု' ဟာ hardware အစိတ်အပိုင်း အမျိုးမျိုး ပေါင်းစပ်ထားတာဖြစ်ပြီး၊ တစ်ခုချင်းစီ ဘာလုပ်လဲဆိုတာ နားလည်ရင် Local AI run တာနဲ့ ပတ်သက်တဲ့ mystery အများကြီး ပျောက်သွားမှာပါ။

CPU (central processing unit) ဟာ ခင်ဗျား computer ရဲ့ general-purpose ဦးနှောက်ပါ - ဘာမဆို လုပ်နိုင်တယ်၊ model run တာအပါအဝင်ပေါ့၊ ဒါပေမယ့် instruction တွေကို တစ်ခုပြီးတစ်ခု အများစု process လုပ်ပါတယ်။

GPU (graphics processing unit) ကတော့ မူလက video game graphics render လုပ်ဖို့ တည်ဆောက်ထားတာဖြစ်ပေမယ့်၊ ဒီအလုပ်ဟာ ရိုးရှင်းတဲ့ math တွေကို data အများကြီးပေါ်မှာ တစ်ပြိုင်နက် လုပ်ဖို့ လိုအပ်တာမို့ - ဒါက neural network run တာအတွက်လည်း အတိအကျ လိုအပ်ချက်ပါပဲ။ ဒါကြောင့် GPU က model ကို CPU ထက် သိသိသာသာ မြန်မြန် run နိုင်တယ်၊ ဒါပေမယ့် CPU ကလည်း ဖြေးဖြေးနဲ့ ရနိုင်ပါသေးတယ်။

device အသစ်တွေမှာ NPU (neural processing unit) လို့ခေါ်တဲ့ ဒီလို AI math အတွက် တိတိကျကျ တည်ဆောက်ထားတဲ့ chip လည်း ပါလာနိုင်ပါတယ်။

Memory ကလည်း tier နှစ်ခုအလားတူပါပဲ။ RAM ဟာ ခင်ဗျား computer ရဲ့ general working memory ဖြစ်ပြီး ဖွင့်ထားတဲ့ program တိုင်း share သုံးပါတယ်။ VRAM ကတော့ GPU ပေါ်မှာ တိုက်ရိုက်နေတဲ့ memory ဖြစ်ပြီး graphics နဲ့ GPU calculation အတွက်ပဲ ချန်ထားတာ၊ RAM ထက် access လုပ်ရတာ ပိုမြန်ပေမယ့် size ပိုသေးငယ်ပါတယ်။

model ကို GPU ပေါ်မှာ run တဲ့အခါ VRAM ထဲကို ဝင်အောင်းရမယ်၊ CPU ပေါ်မှာ run တဲ့အခါ RAM ထဲကို ဝင်အောင်းရမယ်ပါ။ လမ်းကြောင်းနှစ်ခုစလုံး မှားတာမဟုတ်ပါဘူး - VRAM အများကြီးပါတဲ့ GPU က speed ပေးတယ်၊ ဒါပေမယ့် CPU-only laptop တစ်ခုကလည်း model သေးလေးတွေကို comfortable စွာ run နိုင်ပါသေးတယ်၊ အဖြေတစ်ခုစီအတွက် ခဏပိုစောင့်ရမယ်ဆိုတာပဲ ကွာသွားမှာပါ။

CPU
computer ရဲ့ general-purpose processor ဖြစ်ပြီး instruction အများစုကို တစ်ခုပြီးတစ်ခု run ပါတယ်။
GPU
အရင်က graphics အတွက် တည်ဆောက်ခဲ့ပေမယ့် ယခု AI model calculation တွေအတွက်လည်း တစ်ပြိုင်နက် ကောင်းစွာအလုပ်လုပ်တဲ့ chip ပါ။
NPU
AI math အတွက် တိတိကျကျ တည်ဆောက်ထားတဲ့ chip အသစ်တစ်မျိုးဖြစ်ပြီး device အသစ်အချို့မှာ ပါလာတတ်ပါတယ်။
RAM
computer ဖွင့်ထားတဲ့ program အားလုံး share သုံးတဲ့ general working memory ပါ။
VRAM
GPU ပေါ်မှာ တိုက်ရိုက်ရှိတဲ့ memory ဖြစ်ပြီး RAM ထက် မြန်ပေမယ့် size သေးငယ်ပါတယ်။
Storage
model file ကို run ခင် ခေတ္တထားတဲ့ hard drive / SSD ပါ - RAM/VRAM နဲ့ မတူဘဲ run နေချိန်မှာ တိုက်ရိုက်မသုံးပါ။
text
WHERE A MODEL CAN LOAD
----------------------
YOUR DEVICE
-----------
  [CPU] <---> [RAM]
    |            model can load here (slower, always works)
    |
  [GPU] <---> [VRAM]   (optional, if present)
                 model can load here (faster, limited size)

  STORAGE (disk): where the model file sits before loading

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

hobbyist တစ်ယောက်က ရှိပြီးသား laptop ပေါ်မှာ local model run ကြည့်ချင်တယ်ဆိုပါစို့ - GPU အသစ် ဝယ်ဖို့ budget မရှိဘူး။ CPU-only inference ဟာ တကယ်အလုပ်လုပ်တယ်၊ ဖြေးဖြေးပဲဆိုတာ သိထားရင် ဘာမှ အရင်ဝယ်စရာမလိုဘဲ model သေးလေးနဲ့ ဒီနေ့ပဲ စတင်နိုင်ပါတယ်။

ဒါကို gaming PC ပါတဲ့ developer တစ်ယောက်နဲ့ နှိုင်းယှဉ်ကြည့်ပါ - GPU နဲ့ VRAM 12GB ရှိပြီး gaming session ပြင်ပမှာ အများစု idle နေတဲ့ developer အတွက်တော့၊ model မရွေးမီ VRAM size စစ်ကြည့်တာက model တစ်ခုဟာ GPU ပေါ်မှာ မြန်မြန် run မလား ဒါမှမဟုတ် ဖြေးတဲ့ CPU memory ဆီ fallback ဖြစ်မလားဆိုတာ ချက်ချင်းသိစေပါတယ်။

နှစ်ခုစလုံးမှာ လက်တွေ့ skill ကတော့ တူတူပါပဲ - model ဘယ်ဟာမှ မ download လုပ်ခင် ခင်ဗျားမှာ hardware ဘာရှိလဲ ကြည့်ပါ - RAM ဘယ်လောက်ရှိလဲ၊ GPU ရှိလား၊ VRAM ဘယ်လောက်ပါလဲ - ဒါမှ download မအောင်မြင်ပြီးမှ mismatch တွေ့မယ့်အစား ခင်ဗျား စက်နဲ့ ကိုက်ညီတဲ့ model size ရွေးနိုင်မှာပါ။

အတူတူ စမ်းရေးကြည့်မယ်

python
def estimate_memory_gb(params_billion, bytes_per_param):
    """Rough rule of thumb: memory needed = parameter count x bytes
    per parameter. Real usage is a bit higher (activations, overhead)
    but this gives a useful ballpark before downloading a model."""
    total_bytes = params_billion * 1_000_000_000 * bytes_per_param
    return total_bytes / (1024 ** 3)


models = [
    ("TinyLlama 1.1B (FP16)", 1.1, 2),
    ("Llama 7B (FP16)", 7, 2),
    ("Llama 70B (FP16)", 70, 2),
]

for name, params_b, bytes_per_param in models:
    gb = estimate_memory_gb(params_b, bytes_per_param)
    print(f"{name}: ~{gb:.2f} GB")
You should see
model size သုံးခုအတွက် FP16 မှာ လိုအပ်တဲ့ memory (GB) ကို ခန့်မှန်းတွက်ချက်ထားပါတယ်။ Output အတိအကျမှာ - TinyLlama 1.1B (FP16): ~2.05 GB
Llama 7B (FP16): ~13.04 GB
Llama 70B (FP16): ~130.39 GB

၅ မိနစ် စမ်းကြည့်

models list ထဲကို ("Mistral 7B (FP16)", 7, 2) အသစ်တစ်ခု ထည့်ပြီး ပြန် run ကြည့်ပါ။ ပြီးရင် bytes_per_param ကို 4 (FP32) ပြောင်းကြည့်ပြီး result ဘယ်လောက်ကွာသွားလဲ လေ့လာပါ။

သတိလေးတစ်ချက်

'local LLM run ချင်ရင် GPU မရှိမဖြစ်လိုတယ်' လို့ ထင်မှတ်ခြင်း - CPU-only inference ဟာ ဖြေးပေမယ့် တကယ်အလုပ်လုပ်ပါတယ်

RAM နဲ့ VRAM ကို ရောထွေးမိပြီး model ကို VRAM ထက် ကျော်လွန်တဲ့ size ရွေးမိခြင်း

Wikipedia: Graphics processing unitLocal AI / Local LLM

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • 'local LLM run ချင်ရင် GPU မရှိမဖြစ်လိုတယ်' လို့ ထင်မှတ်ခြင်း - CPU-only inference ဟာ ဖြေးပေမယ့် တကယ်အလုပ်လုပ်ပါတယ်
  • RAM နဲ့ VRAM ကို ရောထွေးမိပြီး model ကို VRAM ထက် ကျော်လွန်တဲ့ size ရွေးမိခြင်း
  • Model (သို့) tool အသစ်တစ်ခုကို production/daily-use workflow ထဲ တိုက်ရိုက်မထည့်ခင် သေးငယ်တဲ့ scale နဲ့ အရင်စမ်းကြည့်ပါ။

လေ့ကျင့်ခန်း

models list ထဲကို ("Mistral 7B (FP16)", 7, 2) အသစ်တစ်ခု ထည့်ပြီး ပြန် run ကြည့်ပါ။ ပြီးရင် bytes_per_param ကို 4 (FP32) ပြောင်းကြည့်ပြီး result ဘယ်လောက်ကွာသွားလဲ လေ့လာပါ။

You'll know it worked when: model size သုံးခုအတွက် FP16 မှာ လိုအပ်တဲ့ memory (GB) ကို ခန့်မှန်းတွက်ချက်ထားပါတယ်။ Output အတိအကျမှာ - TinyLlama 1.1B (FP16): ~2.05 GB Llama 7B (FP16): ~13.04 GB Llama 70B (FP16): ~130.39 GB

Hardware အခြေခံများ - CPU, GPU, RAM, VRAM | Thuta Learning