နားလည်ထားရမယ့် အချက်
လူတွေက “local AI” ဆိုရင် app တစ်ခုတည်းကို download လုပ်ပြီး စကားပြောလို့ရတဲ့ software တစ်ခုလို့ ထင်လေ့ရှိကြပါတယ်။ တကယ်တော့ local AI ecosystem ဆိုတာ layer အမျိုးမျိုးနဲ့ ဖွဲ့စည်းထားတာဖြစ်ပြီး၊ layer တစ်ခုစီက ပြဿနာမတူညီတာတွေကို ဖြေရှင်းပေးနေကြပါတယ်။ brand name တွေကို အလွတ်ကျက်တာထက် layer တွေကို နားလည်ဖို့က ပိုအရေးကြီးပါတယ်။
- Model file — အောက်ဆုံးမှာ ရှိပြီး GGUF (သို့) safetensors format ဖြင့် သိမ်းထားတဲ့ weight ဂဏန်းတွေဖြစ်ကာ ကိုယ်တိုင် တိုက်ရိုက် စကားပြောလို့ မရသေးပါ။
- Runtime — weight တွေကို memory ထဲ load လုပ်ပြီး inference တကယ် လုပ်ဆောင်ပေးတဲ့ engine ဖြစ်ပြီး llama.cpp နဲ့ Ollama ရဲ့ internal engine နှစ်ခုစလုံးက runtime တွေပါ။
- API layer — အများအားဖြင့် local HTTP server တစ်ခုဖြစ်ပြီး /api/chat သို့မဟုတ် /v1/chat/completions လို endpoint တွေကို ဖော်ပြထားတာကြောင့် software တခြားတွေက prompt ပို့ပြီး JSON အဖြစ် response ပြန်ရနိုင်ပါတယ်။
- Application layer — LM Studio, Jan လို desktop chat app တွေ၊ LangChain, LlamaIndex လို developer framework တွေ၊ CrewAI, AutoGen လို agent framework တွေ ပါဝင်ပါတယ်။
- Model hub — Hugging Face, Ollama library လိုမျိုး model file ကို ရှာဖွေ download ရယူနိုင်တဲ့ နေရာ။
- Vector database — Chroma, Qdrant လိုမျိုး embedding တွေကို သိမ်းထားတဲ့ database။
ဒီ vertical stack ကို model hub နဲ့ vector database ဆိုတဲ့ category နှစ်ခုက ဖြတ်ကျော်ပါတယ်။ tool တစ်ခုက category ဘယ်ခုမှာ ရှိလဲ သိထားရင် ဘာပြဿနာကို ဖြေရှင်းပေးလဲ သိနိုင်ပြီး model hub ကို inference run ခိုင်းတာ (သို့) chat app ကို software တခြားတွေအတွက် API serve ခိုင်းတာလိုမျိုး မမျှော်လင့်သင့်ပါ။
- Runtime
- Model weight တွေကို memory ထဲ load လုပ်ပြီး output ထုတ်ဖို့ တွက်ချက်မှု တကယ်လုပ်ဆောင်ပေးတဲ့ engine။ llama.cpp နဲ့ Ollama ရဲ့ internal engine နှစ်ခုစလုံးဟာ runtime တွေဖြစ်ကြပါတယ်။
- Inference Server
- runtime တစ်ခုကို network-facing API (အများအားဖြင့် HTTP) နဲ့ ထပ်ပတ်ထားတာပါ၊ command line တစ်ခုတည်းသာမက software တခြားတွေကလည်း request ပို့ပြီး response ရနိုင်အောင် background မှာ အမြဲ run နေတတ်ပါတယ်။
MODEL FILE TO APPLICATION STACK
-------------------------------
------------------------------------------------------------
APPLICATION LM Studio, LangChain, LlamaIndex, CrewAI, AutoGen
^ (chat apps, dev frameworks, agent frameworks)
|
calls
|
API LAYER HTTP server: /api/chat, /v1/chat/completions
^ (JSON request in, JSON response out)
|
exposed by
|
RUNTIME llama.cpp, Ollama engine
^ (loads weights, runs inference)
|
loads
|
MODEL FILE weights on disk: model.gguf, model.safetensors
(just numbers -- cannot be "talked to" directly)
SIDE CATEGORIES (cut across the stack, not the vertical flow):
MODEL HUBS Hugging Face, Ollama Library -> source of files
VECTOR DATABASES Chroma, Qdrant -> store embeddingsလက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်
သင့်team က internal tool တစ်ခုအတွက် “document တွေနဲ့ စကားပြော” feature ထည့်ချင်တယ်လို့ ယူဆပါစို့၊ တစ်ယောက်က planning doc ထဲမှာ name list တစ်ခု ချထားလိုက်တယ် — Ollama, LM Studio, LangChain, Chroma, Hugging Face, CrewAI။ mental map မရှိရင် ဒီ list က ရွေးချယ်ရမယ့် ပြိုင်ဘက် product ခုနစ်ခုလိုပဲ ထင်ရပါတယ်။
- Runtime တစ်ခု — model run ဖို့ (Ollama ကို ရွေးမယ်၊ team က terminal နဲ့ အကျွမ်းတဝင်ရှိပြီး default API လိုချင်လို့)။
- Model hub တစ်ခု — model ရွေးဖို့ (Hugging Face ကို option compare ဖို့ သုံးပေမယ့် နောက်ဆုံး model ကို Ollama library ကနေပဲ pull လုပ်နိုင်ပါတယ်)။
- Vector database တစ်ခု — document embedding တွေ သိမ်းဖို့ (Chroma ကို ရွေးမယ်၊ version ပထမဆုံးအတွက် local run ရလွယ်လို့)။
- Developer framework တစ်ခု (optional) — HTTP call တွေကို လက်နဲ့ ရေးမယ့်အစား component တွေ ချိတ်ဆက်ဖို့ (LangChain ကို သုံးမယ်)။
LM Studio နဲ့ CrewAI တို့ကတော့ list ထဲက ကျန်ခဲ့ပါတယ် — ညံ့လို့မဟုတ်ဘဲ GUI desktop app တစ်ခုနဲ့ multi-agent orchestrator တစ်ခုက ဒီ feature အတွက် အခုချိန်မှာ မလိုအပ်သေးလို့ပါ။ အောက်ပါ exercise မှာ tool list တခြားတစ်ခုနဲ့ ဒီလို classification လုပ်ခိုင်းထားတာကြောင့် code မထိခင် “ဒီ tool က layer ဘယ်ခုမှာ တကယ်ရှိလဲ” လို့ မေးတဲ့အလေ့အထ ကို ရအောင် လုပ်ပါ။
အတူတူ စမ်းရေးကြည့်မယ်
CATEGORY_BY_TOOL = {
"LM Studio": "Desktop App",
"Jan": "Desktop App",
"GPT4All": "Desktop App",
"Ollama": "CLI Runtime",
"llama.cpp": "CLI Runtime",
"vLLM": "Inference Server",
"text-generation-webui": "Inference Server",
"LangChain": "Developer Framework",
"LlamaIndex": "Developer Framework",
"Hugging Face Hub": "Model Hub",
"Ollama Library": "Model Hub",
"Chroma": "Vector Database",
"Qdrant": "Vector Database",
"AutoGen": "Agent Framework",
"CrewAI": "Agent Framework",
}
def classify_tools(tool_names):
result = {}
for name in tool_names:
result[name] = CATEGORY_BY_TOOL.get(name, "Uncategorized")
return result
if __name__ == "__main__":
tools = ["Ollama", "LM Studio", "vLLM", "LangChain", "Chroma", "CrewAI", "Hugging Face Hub"]
classified = classify_tools(tools)
for tool_name, category in classified.items():
print(f"{tool_name}: {category}")
Script ကို run လိုက်ရင် tool name တစ်ခုစီကို category နဲ့ တွဲပြီး တစ်ကြောင်းချင်း print ထုတ်ပါတယ် — Ollama: CLI Runtime, LM Studio: Desktop App, vLLM: Inference Server, LangChain: Developer Framework, Chroma: Vector Database, CrewAI: Agent Framework, Hugging Face Hub: Model Hub ဆိုပြီး အတိအကျ classification အတိုင်း ထွက်ပါတယ်။၅ မိနစ် စမ်းကြည့်
`CATEGORY_BY_TOOL` dictionary ထဲကို tool နောက်ထပ် သုံးခု ထပ်ထည့်ပါ — အခုထိ category မပါသေးတဲ့ tool အစစ်တွေကို ရွေးပါ (ဥပမာ Jan လို desktop app, (သို့) AutoGen လို agent framework)။ ပြီးရင် သင်ထည့်ထားတဲ့ tool တစ်ခုနဲ့ dictionary ထဲ လုံးဝမပါတဲ့ tool တစ်ခု ပါဝင်တဲ့ list ကို `classify_tools` နဲ့ run ကြည့်ပြီး၊ မသိတဲ့ tool က crash မဖြစ်ဘဲ `Uncategorized` လို့ မှန်ကန်စွာ print ထွက်လာမလား စစ်ဆေးပါ။
သတိလေးတစ်ချက်
“local AI” ကို layer အမျိုးမျိုးပါတဲ့ stack အစား download လုပ်နိုင်တဲ့ app တစ်ခုတည်းလို့ ထင်မှတ်ခြင်း။
category တစ်ခုက tool ကို category တခြားရဲ့ အလုပ်ကို လုပ်ခိုင်းမိခြင်း — ဥပမာ model hub က inference run နိုင်တယ်လို့ ထင်ခြင်း (သို့) chat app က software တခြားတွေအတွက် API serve နိုင်တယ်လို့ ယူဆခြင်း။
Ollama on GitHub — Local AI / Local LLM