Thuta Learning
Local AI / Local LLM
IntermediateAIbeginner

Local AI ဆော့ဖ်ဝဲ Ecosystem

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • Local AI ဆော့ဖ်ဝဲ Ecosystem concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • Diagram ကို ဖတ်ပြီး architecture ထဲမှာ data/request ဘယ်လိုစီးဆင်းသလဲ ခြေရာခံနိုင်ရန်
  • ကိုယ့် hardware/use case အတွက် ဘယ်လို ရွေးချယ်သင့်သလဲ ဆုံးဖြတ်နိုင်ရန်

နားလည်ထားရမယ့် အချက်

လူတွေက “local AI” ဆိုရင် app တစ်ခုတည်းကို download လုပ်ပြီး စကားပြောလို့ရတဲ့ software တစ်ခုလို့ ထင်လေ့ရှိကြပါတယ်။ တကယ်တော့ local AI ecosystem ဆိုတာ layer အမျိုးမျိုးနဲ့ ဖွဲ့စည်းထားတာဖြစ်ပြီး၊ layer တစ်ခုစီက ပြဿနာမတူညီတာတွေကို ဖြေရှင်းပေးနေကြပါတယ်။ brand name တွေကို အလွတ်ကျက်တာထက် layer တွေကို နားလည်ဖို့က ပိုအရေးကြီးပါတယ်။

  • Model file — အောက်ဆုံးမှာ ရှိပြီး GGUF (သို့) safetensors format ဖြင့် သိမ်းထားတဲ့ weight ဂဏန်းတွေဖြစ်ကာ ကိုယ်တိုင် တိုက်ရိုက် စကားပြောလို့ မရသေးပါ။
  • Runtime — weight တွေကို memory ထဲ load လုပ်ပြီး inference တကယ် လုပ်ဆောင်ပေးတဲ့ engine ဖြစ်ပြီး llama.cpp နဲ့ Ollama ရဲ့ internal engine နှစ်ခုစလုံးက runtime တွေပါ။
  • API layer — အများအားဖြင့် local HTTP server တစ်ခုဖြစ်ပြီး /api/chat သို့မဟုတ် /v1/chat/completions လို endpoint တွေကို ဖော်ပြထားတာကြောင့် software တခြားတွေက prompt ပို့ပြီး JSON အဖြစ် response ပြန်ရနိုင်ပါတယ်။
  • Application layer — LM Studio, Jan လို desktop chat app တွေ၊ LangChain, LlamaIndex လို developer framework တွေ၊ CrewAI, AutoGen လို agent framework တွေ ပါဝင်ပါတယ်။
  • Model hub — Hugging Face, Ollama library လိုမျိုး model file ကို ရှာဖွေ download ရယူနိုင်တဲ့ နေရာ။
  • Vector database — Chroma, Qdrant လိုမျိုး embedding တွေကို သိမ်းထားတဲ့ database။

ဒီ vertical stack ကို model hub နဲ့ vector database ဆိုတဲ့ category နှစ်ခုက ဖြတ်ကျော်ပါတယ်။ tool တစ်ခုက category ဘယ်ခုမှာ ရှိလဲ သိထားရင် ဘာပြဿနာကို ဖြေရှင်းပေးလဲ သိနိုင်ပြီး model hub ကို inference run ခိုင်းတာ (သို့) chat app ကို software တခြားတွေအတွက် API serve ခိုင်းတာလိုမျိုး မမျှော်လင့်သင့်ပါ။

Runtime
Model weight တွေကို memory ထဲ load လုပ်ပြီး output ထုတ်ဖို့ တွက်ချက်မှု တကယ်လုပ်ဆောင်ပေးတဲ့ engine။ llama.cpp နဲ့ Ollama ရဲ့ internal engine နှစ်ခုစလုံးဟာ runtime တွေဖြစ်ကြပါတယ်။
Inference Server
runtime တစ်ခုကို network-facing API (အများအားဖြင့် HTTP) နဲ့ ထပ်ပတ်ထားတာပါ၊ command line တစ်ခုတည်းသာမက software တခြားတွေကလည်း request ပို့ပြီး response ရနိုင်အောင် background မှာ အမြဲ run နေတတ်ပါတယ်။
text
MODEL FILE TO APPLICATION STACK
-------------------------------
------------------------------------------------------------
APPLICATION   LM Studio, LangChain, LlamaIndex, CrewAI, AutoGen
                 ^  (chat apps, dev frameworks, agent frameworks)
                 |
              calls
                 |
API LAYER     HTTP server: /api/chat, /v1/chat/completions
                 ^  (JSON request in, JSON response out)
                 |
          exposed by
                 |
RUNTIME       llama.cpp, Ollama engine
                 ^  (loads weights, runs inference)
                 |
             loads
                 |
MODEL FILE    weights on disk: model.gguf, model.safetensors
              (just numbers -- cannot be "talked to" directly)

SIDE CATEGORIES (cut across the stack, not the vertical flow):
  MODEL HUBS        Hugging Face, Ollama Library -> source of files
  VECTOR DATABASES  Chroma, Qdrant                -> store embeddings

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

သင့်team က internal tool တစ်ခုအတွက် “document တွေနဲ့ စကားပြော” feature ထည့်ချင်တယ်လို့ ယူဆပါစို့၊ တစ်ယောက်က planning doc ထဲမှာ name list တစ်ခု ချထားလိုက်တယ် — Ollama, LM Studio, LangChain, Chroma, Hugging Face, CrewAI။ mental map မရှိရင် ဒီ list က ရွေးချယ်ရမယ့် ပြိုင်ဘက် product ခုနစ်ခုလိုပဲ ထင်ရပါတယ်။

  • Runtime တစ်ခု — model run ဖို့ (Ollama ကို ရွေးမယ်၊ team က terminal နဲ့ အကျွမ်းတဝင်ရှိပြီး default API လိုချင်လို့)။
  • Model hub တစ်ခု — model ရွေးဖို့ (Hugging Face ကို option compare ဖို့ သုံးပေမယ့် နောက်ဆုံး model ကို Ollama library ကနေပဲ pull လုပ်နိုင်ပါတယ်)။
  • Vector database တစ်ခု — document embedding တွေ သိမ်းဖို့ (Chroma ကို ရွေးမယ်၊ version ပထမဆုံးအတွက် local run ရလွယ်လို့)။
  • Developer framework တစ်ခု (optional) — HTTP call တွေကို လက်နဲ့ ရေးမယ့်အစား component တွေ ချိတ်ဆက်ဖို့ (LangChain ကို သုံးမယ်)။

LM Studio နဲ့ CrewAI တို့ကတော့ list ထဲက ကျန်ခဲ့ပါတယ် — ညံ့လို့မဟုတ်ဘဲ GUI desktop app တစ်ခုနဲ့ multi-agent orchestrator တစ်ခုက ဒီ feature အတွက် အခုချိန်မှာ မလိုအပ်သေးလို့ပါ။ အောက်ပါ exercise မှာ tool list တခြားတစ်ခုနဲ့ ဒီလို classification လုပ်ခိုင်းထားတာကြောင့် code မထိခင် “ဒီ tool က layer ဘယ်ခုမှာ တကယ်ရှိလဲ” လို့ မေးတဲ့အလေ့အထ ကို ရအောင် လုပ်ပါ။

အတူတူ စမ်းရေးကြည့်မယ်

python
CATEGORY_BY_TOOL = {
    "LM Studio": "Desktop App",
    "Jan": "Desktop App",
    "GPT4All": "Desktop App",
    "Ollama": "CLI Runtime",
    "llama.cpp": "CLI Runtime",
    "vLLM": "Inference Server",
    "text-generation-webui": "Inference Server",
    "LangChain": "Developer Framework",
    "LlamaIndex": "Developer Framework",
    "Hugging Face Hub": "Model Hub",
    "Ollama Library": "Model Hub",
    "Chroma": "Vector Database",
    "Qdrant": "Vector Database",
    "AutoGen": "Agent Framework",
    "CrewAI": "Agent Framework",
}


def classify_tools(tool_names):
    result = {}
    for name in tool_names:
        result[name] = CATEGORY_BY_TOOL.get(name, "Uncategorized")
    return result


if __name__ == "__main__":
    tools = ["Ollama", "LM Studio", "vLLM", "LangChain", "Chroma", "CrewAI", "Hugging Face Hub"]
    classified = classify_tools(tools)
    for tool_name, category in classified.items():
        print(f"{tool_name}: {category}")
You should see
Script ကို run လိုက်ရင် tool name တစ်ခုစီကို category နဲ့ တွဲပြီး တစ်ကြောင်းချင်း print ထုတ်ပါတယ် — Ollama: CLI Runtime, LM Studio: Desktop App, vLLM: Inference Server, LangChain: Developer Framework, Chroma: Vector Database, CrewAI: Agent Framework, Hugging Face Hub: Model Hub ဆိုပြီး အတိအကျ classification အတိုင်း ထွက်ပါတယ်။

၅ မိနစ် စမ်းကြည့်

`CATEGORY_BY_TOOL` dictionary ထဲကို tool နောက်ထပ် သုံးခု ထပ်ထည့်ပါ — အခုထိ category မပါသေးတဲ့ tool အစစ်တွေကို ရွေးပါ (ဥပမာ Jan လို desktop app, (သို့) AutoGen လို agent framework)။ ပြီးရင် သင်ထည့်ထားတဲ့ tool တစ်ခုနဲ့ dictionary ထဲ လုံးဝမပါတဲ့ tool တစ်ခု ပါဝင်တဲ့ list ကို `classify_tools` နဲ့ run ကြည့်ပြီး၊ မသိတဲ့ tool က crash မဖြစ်ဘဲ `Uncategorized` လို့ မှန်ကန်စွာ print ထွက်လာမလား စစ်ဆေးပါ။

သတိလေးတစ်ချက်

“local AI” ကို layer အမျိုးမျိုးပါတဲ့ stack အစား download လုပ်နိုင်တဲ့ app တစ်ခုတည်းလို့ ထင်မှတ်ခြင်း။

category တစ်ခုက tool ကို category တခြားရဲ့ အလုပ်ကို လုပ်ခိုင်းမိခြင်း — ဥပမာ model hub က inference run နိုင်တယ်လို့ ထင်ခြင်း (သို့) chat app က software တခြားတွေအတွက် API serve နိုင်တယ်လို့ ယူဆခြင်း။

Ollama on GitHubLocal AI / Local LLM

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • “local AI” ကို layer အမျိုးမျိုးပါတဲ့ stack အစား download လုပ်နိုင်တဲ့ app တစ်ခုတည်းလို့ ထင်မှတ်ခြင်း။
  • category တစ်ခုက tool ကို category တခြားရဲ့ အလုပ်ကို လုပ်ခိုင်းမိခြင်း — ဥပမာ model hub က inference run နိုင်တယ်လို့ ထင်ခြင်း (သို့) chat app က software တခြားတွေအတွက် API serve နိုင်တယ်လို့ ယူဆခြင်း။
  • Model (သို့) tool အသစ်တစ်ခုကို production/daily-use workflow ထဲ တိုက်ရိုက်မထည့်ခင် သေးငယ်တဲ့ scale နဲ့ အရင်စမ်းကြည့်ပါ။

လေ့ကျင့်ခန်း

`CATEGORY_BY_TOOL` dictionary ထဲကို tool နောက်ထပ် သုံးခု ထပ်ထည့်ပါ — အခုထိ category မပါသေးတဲ့ tool အစစ်တွေကို ရွေးပါ (ဥပမာ Jan လို desktop app, (သို့) AutoGen လို agent framework)။ ပြီးရင် သင်ထည့်ထားတဲ့ tool တစ်ခုနဲ့ dictionary ထဲ လုံးဝမပါတဲ့ tool တစ်ခု ပါဝင်တဲ့ list ကို `classify_tools` နဲ့ run ကြည့်ပြီး၊ မသိတဲ့ tool က crash မဖြစ်ဘဲ `Uncategorized` လို့ မှန်ကန်စွာ print ထွက်လာမလား စစ်ဆေးပါ။

You'll know it worked when: Script ကို run လိုက်ရင် tool name တစ်ခုစီကို category နဲ့ တွဲပြီး တစ်ကြောင်းချင်း print ထုတ်ပါတယ် — Ollama: CLI Runtime, LM Studio: Desktop App, vLLM: Inference Server, LangChain: Developer Framework, Chroma: Vector Database, CrewAI: Agent Framework, Hugging Face Hub: Model Hub ဆိုပြီး အတိအကျ classification အတိုင်း ထွက်ပါတယ်။

Local AI ဆော့ဖ်ဝဲ Ecosystem | Thuta Learning