Thuta Learning
Local AI / Local LLM
IntermediateAIbeginner

LM Studio နဲ့ llama.cpp

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • LM Studio နဲ့ llama.cpp concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • Diagram ကို ဖတ်ပြီး architecture ထဲမှာ data/request ဘယ်လိုစီးဆင်းသလဲ ခြေရာခံနိုင်ရန်
  • ကိုယ့် hardware/use case အတွက် ဘယ်လို ရွေးချယ်သင့်သလဲ ဆုံးဖြတ်နိုင်ရန်

နားလည်ထားရမယ့် အချက်

LM Studio နဲ့ llama.cpp ကို အတူတူ ပြောလေ့ရှိပေမယ့် layer မတူဘဲ user အမျိုးအစား မတူတာကို ဝန်ဆောင်ပေးကြပါတယ်၊ ဒါပေမယ့် နောက်ကွယ်မှာတော့ GGUF model file တစ်ခုတည်းကို run နေကြတာ များပါတယ်။

Toolဘာလဲ / ဘယ်သူအတွက်လဲ
llama.cppC/C++ library နဲ့ command-line program အစုအဝေးတစ်ခုဖြစ်ပြီး datacenter GPU မလိုဘဲ CPU ပုံမှန်ပေါ်မှာတောင် language model ကို ထိရောက်စွာ run နိုင်စေပါတယ်။ GGUF model format ကို ရေပန်းစားစေခဲ့တဲ့ low-level engine ဖြစ်ပြီး CPU-only inference, GPU ရှိရင် layer တွေကို GPU ဆီ တစ်ခြမ်း (သို့) အပြည့် offload လုပ်တာ, ပြီးတော့ built-in server mode (HTTP API) ကို support လုပ်ပါတယ်။ terminal နဲ့ အကျွမ်းတဝင်ရှိတဲ့, memory/performance setting ကို ထိန်းချုပ်ချင်တဲ့, local inference ကို ကိုယ်ပိုင် app ထဲ embed ချင်တဲ့ developer တွေအတွက် အကျိုးရှိပါတယ်။
LM Studioengine အမျိုးအစားတူတူပေါ်မှာ တည်ဆောက်ထားတဲ့ graphical desktop app တစ်ခုဖြစ်ပြီး flag နဲ့ config file အစား window နဲ့ button တွေနဲ့ model ရှာ, download, chat လုပ်ချင်တဲ့ user တွေအတွက် ရည်ရွယ်ပါတယ်။ searchable model catalog, click အနည်းငယ်နဲ့ load, ရင်းနှီးတဲ့ chat-window interface, ပြီးတော့ terminal လက်တစ်ချက်မှ မထိဘဲ software တခြားတွေအတွက် OpenAI-compatible API ဖော်ပြပေးတဲ့ local server mode ကို ဖွင့်နိုင်ပါတယ်။

tool ဘယ်ခုမှ ပိုကောင်းတယ်လို့ တင်းကျပ်စွာ မပြောနိုင်ပါ — visual learner (သို့) မြန်မြန် prototype လုပ်ချင်သူက LM Studio ရဲ့ GUI ကနေ အကျိုးရှိပြီး, workflow automate လုပ်ချင် (သို့) performance အမြင့်ဆုံးယူချင်တဲ့ developer က llama.cpp ရဲ့ တိုက်ရိုက် ထိန်းချုပ်မှုကနေ အကျိုးရှိပါတယ်၊ လူအများစုကတော့ task မတူညီရင် နှစ်ခုစလုံးကို သုံးကြပါတယ်။

text
TWO PATHS TO A RUNNING MODEL
----------------------------
------------------------------------------------------------
                    GGUF MODEL FILE
                    /              \\
                   /                \\
        GUI PATH                       CLI / LIBRARY PATH
       LM Studio                          llama.cpp
      search model                    build or download
      click "Download"                point at model path
      click "Load"                    run server binary
      chat in a window                call its HTTP API
                   \\                /
                    \\              /
                     v            v
              A LOCAL MODEL, RUNNING
           (answers prompts, either way)

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

target တူတူရှိတဲ့ learner နှစ်ယောက်ကို ပုံဖော်ကြည့်ပါ — GGUF model တစ်ခုကို local run ပြီး စမ်းသပ်ကြည့်ချင်ကြတယ်ဆိုပါစို့။

command line နဲ့ မရင်းနှီးသေးတဲ့ ပထမလူက LM Studio ကို download လုပ်, model search tab ဖွင့်, model name ရိုက်, download click, progress bar စောင့်, ပြီးရင် chat tab ဖွင့်ပြီး စာစတင်ရိုက်ပါတယ်; လမ်းကြောင်းတစ်ခုလုံး GUI ထဲကနေ တစ်ခါမှ မထွက်ပါဘူး။

တည်ဆောက်ရတာ ရင်းနှီးတဲ့ ဒုတိယလူကတော့ llama.cpp release ကို compile (သို့) download လုပ်ပြီး disk ပေါ်က GGUF file တစ်ခုနဲ့ port ကို ညွှန်းပြီး server binary ကို run ပါတယ်၊ ခန့်မှန်းချက်ပုံစံက `<binary> --model <path-to-model.gguf> --port <port>` လိုမျိုးဖြစ်ပြီး local HTTP endpoint တစ်ခုကို ချက်ချင်း script ရေးနိုင်အောင် ရရှိပါတယ်။ နှစ်ယောက်စလုံးဟာ ဒီသင်ခန်းစာ diagram ထဲက နေရာတစ်ခုတည်းကို ရောက်သွားကြပါတယ် — local model တစ်ခု, run နေတယ်, prompt တွေကို ဖြေနေတယ်။

ကွာခြားတာကတော့ နောက်ဆက်တွဲမှာပါ။ LM Studio user က code ကနေ ခေါ်ချင်လာရင် reinstall မလိုဘဲ toggle တစ်ခုနဲ့ app ရဲ့ built-in server mode ကို ဖွင့်လို့ရပါတယ်။ llama.cpp user ကတော့ server ရှိပြီးသားဖြစ်လို့ GUI ဘယ်တော့မှ မဖွင့်ဘဲ script, cron job, program တခြားထဲကို ချိတ်ဆက်လို့ ရပါတယ်။ လမ်းကြောင်းနှစ်ခုစလုံးက ဘယ်တစ်ခုမှ ပို “တကယ်” မဟုတ်ပါဘူး — ဒီသင်ခန်းစာ ရည်ရွယ်ထားတဲ့ on-ramp နှစ်ခုပါပဲ။

အတူတူ စမ်းရေးကြည့်မယ်

bash
# CPU-only, run the server on the default port
./llama-server --model ./models/llama-3.2.gguf --port 8080

# offload some layers to a GPU if one is available
./llama-server --model ./models/llama-3.2.gguf --port 8080 --n-gpu-layers 20

# one-shot CLI generation instead of a server
./llama-cli --model ./models/llama-3.2.gguf -p "Explain what a runtime is."
You should see
command ပထမနှစ်ခုက `llama-server` ကို run ခိုင်းပါတယ်၊ ဒါက ပေးထားတဲ့ GGUF model ကို memory ထဲ load လုပ်ပြီး ပေးထားတဲ့ port ပေါ်မှာ HTTP request တွေကို နားထောင်စပါတယ် — `--n-gpu-layers` ထားရင် model ရဲ့ layer အချို့ကို CPU ပေါ်မှာ အကုန်run မယ့်အစား ရနိုင်တဲ့ GPU ဆီ offload လုပ်ပေးလို့ generation ကို ပုံမှန် မြန်စေပါတယ်။ command နှစ်ခုစလုံးက chat reply ကို ကိုယ်တိုင် print မထုတ်ပါဘူး; server က request တွေကို စောင့်နေရုံပါပဲ၊ နောက်သင်ခန်းစာတွေက API pattern နဲ့ ဆင်တူပါတယ်။ တတိယ command ကတော့ `llama-cli` ကို one-shot mode နဲ့ run ခိုင်းတာဖြစ်ပြီး model ကို load လုပ်, ပေးထားတဲ့ prompt အတွက် completion ကို terminal ထဲမှာ တိုက်ရိုက် generate လုပ်ပြီး ပြီးဆုံးသွားပါတယ်၊ server မပါဝင်ပါဘူး။ flag name အတိအကျဟာ llama.cpp release အသစ်တွေနဲ့ ပြောင်းလဲနိုင်လို့ flag တစ်ခုကို အားကိုးမသုံးခင် install ထားတဲ့ binary ပေါ်မှာ `--help` ကို အမြဲ စစ်ဆေးပါ။

၅ မိနစ် စမ်းကြည့်

(llama.cpp ကို local မှာ build ထားပြီးသားမဟုတ်ရင်) run ကြည့်ဖို့ မလိုပါဘူး — `./models/mistral-7b.gguf` ဆိုတဲ့ hypothetical model ကို port `9090` ပေါ်မှာ layer ၁၅ခု GPU ဆီ offload လုပ်ပြီး `llama-server` စတင်ဖို့ command ကို ရေးကြည့်ပါ။ ပြီးရင် ရလဒ်တူတူ ရောက်အောင် (ဒီ model အတွက် local ကနေ ရောက်နိုင်တဲ့ chat API တစ်ခု) LM Studio နဲ့ ဘယ်လို step တွေ (search, load, server mode toggle, port ဘယ်လောက်) လုပ်ရမလဲ ရိုးရှင်းတဲ့ စာလုံးနဲ့ ရေးကြည့်ပါ။

သတိလေးတစ်ချက်

llama.cpp ရဲ့ CLI flag အတိအကျက release အသစ်တွေနဲ့ ဘယ်တော့မှ မပြောင်းဘူးလို့ ယူဆခြင်း — tutorial အဟောင်းက flag ကို copy မယူဘဲ install ထားတဲ့ binary ပေါ်မှာ `--help` ကို အမြဲ စစ်ဆေးပါ။

“ဘယ်ဟာက ပိုကောင်းလဲ” ဆိုပြီး LM Studio (သို့) llama.cpp ကို ရွေးခြင်း — GUI အဆင်ပြေမှု နဲ့ script ရေးလို့ ထိန်းချုပ်နိုင်မှု ဘယ်ဟာက workflow အစစ်နဲ့ ကိုက်ညီလဲ ကြည့်သင့်ပါတယ်။

llama.cpp on GitHubLocal AI / Local LLM

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • llama.cpp ရဲ့ CLI flag အတိအကျက release အသစ်တွေနဲ့ ဘယ်တော့မှ မပြောင်းဘူးလို့ ယူဆခြင်း — tutorial အဟောင်းက flag ကို copy မယူဘဲ install ထားတဲ့ binary ပေါ်မှာ `--help` ကို အမြဲ စစ်ဆေးပါ။
  • “ဘယ်ဟာက ပိုကောင်းလဲ” ဆိုပြီး LM Studio (သို့) llama.cpp ကို ရွေးခြင်း — GUI အဆင်ပြေမှု နဲ့ script ရေးလို့ ထိန်းချုပ်နိုင်မှု ဘယ်ဟာက workflow အစစ်နဲ့ ကိုက်ညီလဲ ကြည့်သင့်ပါတယ်။
  • Model (သို့) tool အသစ်တစ်ခုကို production/daily-use workflow ထဲ တိုက်ရိုက်မထည့်ခင် သေးငယ်တဲ့ scale နဲ့ အရင်စမ်းကြည့်ပါ။

လေ့ကျင့်ခန်း

(llama.cpp ကို local မှာ build ထားပြီးသားမဟုတ်ရင်) run ကြည့်ဖို့ မလိုပါဘူး — `./models/mistral-7b.gguf` ဆိုတဲ့ hypothetical model ကို port `9090` ပေါ်မှာ layer ၁၅ခု GPU ဆီ offload လုပ်ပြီး `llama-server` စတင်ဖို့ command ကို ရေးကြည့်ပါ။ ပြီးရင် ရလဒ်တူတူ ရောက်အောင် (ဒီ model အတွက် local ကနေ ရောက်နိုင်တဲ့ chat API တစ်ခု) LM Studio နဲ့ ဘယ်လို step တွေ (search, load, server mode toggle, port ဘယ်လောက်) လုပ်ရမလဲ ရိုးရှင်းတဲ့ စာလုံးနဲ့ ရေးကြည့်ပါ။

You'll know it worked when: command ပထမနှစ်ခုက `llama-server` ကို run ခိုင်းပါတယ်၊ ဒါက ပေးထားတဲ့ GGUF model ကို memory ထဲ load လုပ်ပြီး ပေးထားတဲ့ port ပေါ်မှာ HTTP request တွေကို နားထောင်စပါတယ် — `--n-gpu-layers` ထားရင် model ရဲ့ layer အချို့ကို CPU ပေါ်မှာ အကုန်run မယ့်အစား ရနိုင်တဲ့ GPU ဆီ offload လုပ်ပေးလို့ generation ကို ပုံမှန် မြန်စေပါတယ်။ command နှစ်ခုစလုံးက chat reply ကို ကိုယ်တိုင် print မထုတ်ပါဘူး; server က request တွေကို စောင့်နေရုံပါပဲ၊ နောက်သင်ခန်းစာတွေက API pattern နဲ့ ဆင်တူပါတယ်။ တတိယ command ကတော့ `llama-cli` ကို one-shot mode နဲ့ run ခိုင်းတာဖြစ်ပြီး model ကို load လုပ်, ပေးထားတဲ့ prompt အတွက် completion ကို terminal ထဲမှာ တိုက်ရိုက် generate လုပ်ပြီး ပြီးဆုံးသွားပါတယ်၊ server မပါဝင်ပါဘူး။ flag name အတိအကျဟာ llama.cpp release အသစ်တွေနဲ့ ပြောင်းလဲနိုင်လို့ flag တစ်ခုကို အားကိုးမသုံးခင် install ထားတဲ့ binary ပေါ်မှာ `--help` ကို အမြဲ စစ်ဆေးပါ။

LM Studio နဲ့ llama.cpp | Thuta Learning