နားလည်ထားရမယ့် အချက်
ChatGPT လိုမျိုး chatbot တစ်ခုကို စာရိုက်ပြီး မေးတိုင်း၊ မျက်စိမမြင်ရတဲ့ လုပ်ငန်းစဉ်တစ်ခု နောက်ကွယ်မှာ ဖြစ်နေတယ်။ ခင်ဗျားရဲ့ စာသားက internet ကနေတဆင့် အခြားနေရာက strong computer တစ်လုံးဆီ ရောက်သွားပြီး၊ process လုပ်ပြီးမှ ပြန်လည်ရောက်ရှိလာတာပါ။ device → internet → cloud server → model → ပြန်လာတဲ့ ဒီ round trip တစ်ခုလုံးဟာ Cloud AI အလုပ်လုပ်ပုံပါပဲ။
Local AI ကတော့ အလယ်က step ကို လုံးဝ ဖယ်ရှားလိုက်တာပါ။ မေးခွန်းကို ဘယ်နေရာကိုမှ မပို့ဘဲ၊ AI model ကိုယ်တိုင်ကို ခင်ဗျားရဲ့ ကိုယ်ပိုင် computer (သို့) ဖုန်းပေါ်မှာ file တစ်ခုအနေနဲ့ ထားရှိပြီး၊ လုပ်ငန်းစဉ်တစ်ခုလုံးကို ခင်ဗျားပိုင်ဆိုင်တဲ့ hardware ပေါ်မှာပဲ internet လိုအပ်ချက် လုံးဝမရှိဘဲ လုပ်ဆောင်တာပါ။
ဒါက "AI သုံးတယ်" ဆိုတဲ့ အဓိပ္ပာယ်ကို ပြောင်းလဲပစ်တယ်။ model ဆိုတာ mysterious ဖြစ်တဲ့ cloud service မဟုတ်ဘူး။ ဂဏန်းတွေ အပြည့်ပါတဲ့ ဖိုင်ကြီး တစ်ခုပါ၊ learn ထားတဲ့ pattern တွေအပြည့်ပါတဲ့ အလွန်ကြီးမားပြီး အသေးစိတ်ဆုံး ဟင်းချက်စာအုပ်ကြီးတစ်အုပ်နဲ့ ဆင်တူတယ်လို့ ခင်ဗျား မြင်ကြည့်နိုင်ပါတယ်။ ကိုယ်ပိုင် computer က ဒီ file ကို memory ထဲကို load လုပ်ပြီး တွက်ချက်မှုတွေ လုပ်ဆောင်ပါတယ်။
load လုပ်ထားတဲ့ model ကို မေးခွန်းမေးပြီး အဖြေရယူတဲ့ လုပ်ငန်းစဉ်ကို inference လို့ ခေါ်ပါတယ်။ inference ကို local မှာလုပ်တဲ့အခါ ခင်ဗျားရဲ့ မေးခွန်း၊ document၊ (သို့) model ရဲ့ အဖြေဟာ device ထဲကနေ တစ်ခါမှ မထွက်ပါဘူး - အောက်ပါအရာတွေ လုပ်နိုင်မယ့် server တစ်ခုမှ အလယ်မှာ မရှိပါဘူး -
- ခင်ဗျားမေးခဲ့တာကို log မှတ်ခြင်း
- data ကို analyze လုပ်ခြင်း
- data ကို ရောင်းချခြင်း
- data breach ထဲမှာ ပျောက်ဆုံးစေခြင်း
နှိုင်းယှဉ်ရရင် - နိုင်ငံခြားက ဘာသာပြန်ဆရာဆီကို စာပို့ခြင်းနဲ့၊ ခင်ဗျားနား ထိုင်နေတဲ့ နှစ်ဘာသာတတ် သူငယ်ချင်းတစ်ယောက် ချက်ချင်း ဘာသာပြန်ပေးတာနဲ့ ခြားနားချက်ပါပဲ - ဘယ်သူ့ကိုမှ ချရေးမပေးဘဲနဲ့ပေါ့။ Local AI ဆိုတာ အခန်းထဲက သူငယ်ချင်းပါပဲ၊ ဒီသင်ခန်းစာက ဒီသူငယ်ချင်းကို ဘာနဲ့ ပြုလုပ်ထားလဲဆိုတာ တိတိကျကျ နားလည်ဖို့ပါ။
- LLM
- စာသားအများကြီးနဲ့ train လုပ်ထားပြီး prompt တစ်ခုကို ဆက်တိုက် စာသားအဖြစ် ပြန်ဖြေနိုင်တဲ့ AI model အမျိုးအစားပါ (Large Language Model ၏ အတိုကောက်)။
- Model
- AI system ရဲ့ trained ဖြစ်ပြီးသား weights အားလုံးကို သိမ်းထားတဲ့ file တစ်ခုပါ — run လုပ်ဖို့ လိုအပ်တဲ့ အရာအားလုံး ဒီ file ထဲမှာ ပါဝင်ပါတယ်။
- Weights
- Model training အတွင်း learn လုပ်ထားတဲ့ ဂဏန်းအများကြီးပါ — model ရဲ့ 'သိထားတဲ့ အရာ' က ဒီ ဂဏန်းတွေထဲမှာ သိမ်းထားတာပါ။
- Parameters
- Model တစ်ခုထဲက weight အရေအတွက်ပါ — 7B ဆိုရင် weight ၇ ဘီလီယံ ရှိတယ်လို့ ဆိုလိုပြီး၊ model ရဲ့ size ကို ကြမ်းအားဖြင့် ညွှန်ပြပါတယ်။
- Inference
- Train ဖြစ်ပြီးသား model ကို input ပေးပြီး output ထုတ်ခိုင်းတဲ့ process ပါ — 'run လုပ်ခြင်း' လို့ ဆိုနိုင်ပါတယ်။
- Prompt
- Model ဆီ ပို့တဲ့ input စာသားပါ — question, instruction, (သို့) conversation history ဖြစ်နိုင်ပါတယ်။
- Local AI
- AI model ကို cloud server ပေါ်မဟုတ်ဘဲ ကိုယ့် device ပေါ်မှာ တိုက်ရိုက် run လုပ်ခြင်းပါ။
CLOUD AI VS LOCAL AI
--------------------
CLOUD AI
--------
[Your Device] --> [Internet] --> [Cloud Server]
|
[Model]
|
[Your Device] <-- [Internet] <-- [Cloud Server]
data leaves your device and crosses the internet twice
LOCAL AI
--------
[Your Device]
|
[Local Model] (a file loaded into memory)
|
[Answer]
data never leaves your device, no internet hop at allလက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်
ရန်ကုန်က စီးပွားရေးလုပ်ငန်းငယ်ရှင်တစ်ယောက်ကို စဉ်းစားကြည့်ပါ - customer message တွေကို reply ရေးပေးဖို့ AI assistant တစ်ခု လိုချင်ပေမယ့် message အများစုမှာ ဖုန်းနံပါတ်၊ လိပ်စာ၊ order details တွေ ပါဝင်နေလို့ နိုင်ငံခြား company ရဲ့ server ဆီကို ပို့ဖို့ စိတ်မကောင်းဖြစ်နေတဲ့ အခြေအနေပါ။
(သို့) signal လုံးဝမရှိတဲ့ ကားရှည် ခရီးပေါ်ရောက်နေတဲ့ developer တစ်ယောက်က chatbot idea တစ်ခု စမ်းသပ်ချင်နေတာကို စဉ်းစားကြည့်ပါ။ နှစ်ယောက်စလုံးက အခြေခံ လိုအပ်ချက် တူတူပါပဲ - data ကို (သို့) internet bill ကို တခြားသူဆီ ပေးအပ်စရာမလိုဘဲ run နိုင်တဲ့ AI တစ်ခုပါ။ ဒါဟာ Local AI ဖြည့်ဆည်း ပေးတဲ့ ကွက်လပ်အတိအကျပါပဲ။
model ကို ခင်ဗျားထိန်းချုပ်ထားတဲ့ hardware ပေါ်မှာ software အနေနဲ့ run တာဖြစ်လို့ အောက်ပါ အားသာချက်တွေ ရရှိပါတယ် -
- monthly API bill မရှိခြင်း
- third party online ဆက်ရှိနေဖို့ မှီခိုစရာမလိုခြင်း
- data ဘယ်ကိုသွားလဲ ဆိုတဲ့ မေးခွန်း မရှိခြင်း - အခန်းထဲကနေ တစ်ခါမှ မထွက်ခြင်း
ဒီ tutorial ရဲ့ နောက်သင်ခန်းစာတွေမှာ ဒီလို model ကို ဘယ်လို download လုပ်ပြီး run ရမလဲဆိုတာ ပြသွားပါမယ်၊ ဒီသင်ခန်းစာရဲ့ code ကတော့ real model မထိခင် input ဝင်လာတယ် → processing ဒီနေရာမှာပဲဖြစ်တယ် → output ထွက်လာတယ်ဆိုတဲ့ idea ရဲ့ ပုံသဏ္ဍာန်ကို ပြသတဲ့ ရိုးရှင်းအောင် လုပ်ထားတဲ့ နမူနာသေးလေးတစ်ခုသာ ဖြစ်ပါတယ်။
အတူတူ စမ်းရေးကြည့်မယ်
def toy_local_model(user_input):
"""A tiny rule-based stand-in for a real language model.
Real models use learned weights; this uses simple keyword rules
just to show the input -> local processing -> output shape."""
text = user_input.lower()
if "hello" in text or "hi" in text:
return "Hello! I am a tiny local model running on your device."
elif "capital" in text and "france" in text:
return "Paris is the capital of France."
elif "weather" in text:
return "I cannot check the weather - I have no internet connection."
else:
return "I do not know that one yet. I am just a small demo model."
inputs = [
"Hello there!",
"What is the capital of France?",
"What is the weather today?",
"Tell me a joke",
]
for user_input in inputs:
response = toy_local_model(user_input)
print(f"Input: {user_input}")
print(f"Output: {response}")
print("---")
Input/Output ၄ ခု ပြသထားပြီး၊ 'hello' ပါလျှင် greeting ပြန်ပေးမည်၊ 'capital'+'france' ပါလျှင် Paris ဟု ဖြေပေးမည်၊ 'weather' ပါလျှင် internet မရှိကြောင်း ပြောမည်၊ ကျန်တာများအတွက် 'မသိသေးပါ' ဟု ပြန်ဖြေမည်။ Output အတိအကျမှာ - Input: Hello there!
Output: Hello! I am a tiny local model running on your device.
---
Input: What is the capital of France?
Output: Paris is the capital of France.
---
Input: What is the weather today?
Output: I cannot check the weather - I have no internet connection.
---
Input: Tell me a joke
Output: I do not know that one yet. I am just a small demo model.
---၅ မိနစ် စမ်းကြည့်
toy_local_model function ထဲကို "thank you" ကို match လုပ်တဲ့ elif condition အသစ်တစ်ခု ထည့်ပြီး "You are welcome!" လို့ ပြန်ဖြေအောင် ပြင်ကြည့်ပါ။ ပြီးရင် input list ထဲကို "Thanks!" ထည့်ပြီး ပြန် run ကြည့်ပါ။
သတိလေးတစ်ချက်
Local AI ဆိုတာ 'internet လိုအပ်ချက်မရှိဘူး' ဆိုတာနဲ့ 'setup မလိုအပ်ဘူး' ကို ရောထွေးမိတတ်ခြင်း - model ကို ပထမဆုံးအကြိမ် download လုပ်ဖို့တော့ internet လိုအပ်ပါသေးတယ်
'model' နဲ့ 'app' ကို တစ်ခုတည်းလို့ ထင်မှတ်တတ်ခြင်း - model က data file တစ်ခုသာဖြစ်ပြီး၊ ၎င်းကို run ပေးဖို့ runtime software တစ်ခု ခွဲပြီးလိုအပ်ပါသေးတယ်
Wikipedia: Inference engine — Local AI / Local LLM