Thuta Learning
CI/CD with GitHub Actions
AdvancedDevOps & Toolsbeginner

DORA နဲ့ Pipeline ကျန်းမာရေးကို တိုင်းတာခြင်း

ဒီခန်းပြီးရင် ဘာတတ်သွားမလဲ

  • DORA နဲ့ Pipeline ကျန်းမာရေးကို တိုင်းတာခြင်း concept ကို နားလည်ရှင်းပြနိုင်ရန်
  • Diagram ကို ဖတ်ပြီး pipeline ထဲမှာ job တွေ ဘယ်လိုစီးဆင်းပြီး ဘာက gate လုပ်သလဲ ခြေရာခံနိုင်ရန်
  • Workflow file ကို ဖတ်ပြီး ဘယ် job တွေ ဘယ်အစဉ်လိုက် run မလဲ ကြိုတင်ပြောနိုင်ရန်

နားလည်ထားရမယ့် အချက်

DORA ရဲ့ metric လေးခုဟာ စွမ်းဆောင်ရည်မြင့်တဲ့ delivery အသင်းတွေကို ဘာက ခွဲခြားပေးလဲ ဆိုတာ နှစ်ပေါင်းများစွာ သုတေသနလုပ်ထားရာက ထွက်လာတာပါ၊ သူတို့ရဲ့ တန်ဖိုးက လှုပ်ရှားမှုကို မတိုင်းဘဲ ရလဒ်ကို တိုင်းတာလို့ပါ။ deployment frequency — production ဆီ ဘယ်လောက် မကြာခဏ အောင်မြင်စွာ release လုပ်လဲ။ lead time for changes — commit ကနေ production မှာ တကယ် run နေတဲ့အထိ ဘယ်လောက် ကြာလဲ။ change failure rate — deployment တွေထဲက ဘယ်နှစ်ရာခိုင်နှုန်းက ပြန်ပြင်ရတဲ့ ချွတ်ယွင်းချက် ဖြစ်စေလဲ။ failed deployment recovery time — အဲဒီလို ဖြစ်တဲ့အခါ service ပြန်ကောင်းဖို့ ဘယ်လောက် ကြာလဲ။

ပထမနှစ်ခုက throughput ကို၊ နောက်နှစ်ခုက stability ကို ဖော်ပြတယ်။ လူတွေ အံ့သြတတ်တဲ့ တွေ့ရှိချက်က ဒါတွေဟာ အပြန်အလှန် အလဲအလှယ် မဟုတ်ဘူး ဆိုတာပါ — ပိုမကြာခဏ deploy လုပ်တဲ့ အသင်းတွေက ပိုနည်းနည်းပဲ ကျတယ်၊ ပြောင်းလဲမှု သေးသေးလေးတွေက review လုပ်ရ၊ စစ်ရ၊ ပြန်ရုပ်ရ ပိုလွယ်လို့ပါ။ နှေးတာက သတိကြီးတာ မဟုတ်ဘူး။

ဒါကြောင့်လည်း တစ်ခုချင်းစီကို သီးသန့် ဖတ်ရင် ကစားလို့ရပြီး မကြာခဏ အန္တရာယ်ရှိတယ်။ deployment frequency တစ်ခုတည်းကို optimize လုပ်ရင် ဂရုမစိုက်တဲ့ release တွေနဲ့ တက်လာတဲ့ failure rate ကို ရမယ်။ change failure rate တစ်ခုတည်းကို optimize လုပ်ရင် အလုံခြုံဆုံး နည်းလမ်းက လုံးဝ မ deploy တာပဲ — delivery ရပ်သွားပြီးသား အသင်းအတွက် ရမှတ်ပြည့်ပါ။ lead time တစ်ခုတည်းက မြန်မြန် merge လုပ်တာကို ဆုချပြီး တကယ် user ဆီ ရောက်မရောက် ဂရုမစိုက်ဘူး။ speed metric တစ်ခုကို stability metric တစ်ခုနဲ့ အမြဲ ယှဉ်ဖတ်ပါ — တကယ့် တိုးတက်မှုက တစ်ခုကို တိုးစေပြီး ကျန်တစ်ခုကို မဆုတ်ယုတ်စေဘူး။

လက်တွေ့ သတိပေးချက် နှစ်ခု။ အဓိပ္ပာယ်သတ်မှတ်ချက်တွေကို ရေးမှတ်ပြီး မပြောင်းပါနဲ့ — "deployment တစ်ခု" ဆိုတာ ဇန်နဝါရီမှာရော ဇွန်မှာရော အတူတူ ဖြစ်နေရမယ်၊ မဟုတ်ရင် ကိုယ့် trend က ကိုယ့်ရဲ့ အဓိပ္ပာယ်သတ်မှတ်ချက် ပြောင်းလဲမှုကိုပဲ တိုင်းနေတာပါ။ ပြီးတော့ ဒီကိန်းဂဏန်းတွေကို အသင်းအဆင့် တိုးတက်မှု signal အဖြစ်သာ ထားပါ၊ တစ်ဦးချင်း စွမ်းဆောင်ရည် ပစ်မှတ် အဖြစ် ဘယ်တော့မှ မထားပါနဲ့ — အကဲဖြတ်ဖို့ သုံးတဲ့ metric ကို လူတွေက အလွယ်ဆုံး ဖောင်းပွစေနိုင်တဲ့ နည်းနဲ့ ပြည့်မီအောင် လုပ်ကြပါလိမ့်မယ်။

pipeline duration နဲ့ flake rate ကိုလည်း ဒီလေးခုနဲ့ တန်းတူ ပထမတန်းစား သတ်မှတ်ပါ။ သူတို့က ကြိုတင်ညွှန်းကိန်းတွေပါ — lead time နဲ့ change failure rate က မြင်သာအောင် ဆိုးလာတာထက် ရက်သတ္တပတ်များစွာ စောပြီး သူတို့က ဆုတ်ယုတ်နေပါလိမ့်မယ်။

text
THE FOUR METRICS: SPEED VERSUS STABILITY
----------------------------------------
  STABILITY
      ^
      |  change failure rate (lower is better)
      |  failed deployment recovery time (lower is better)
      |
      |            +------------------------------+
      |            | high performers sit up here:  |
      |            | frequent small deploys AND    |
      |            | fast, reliable recovery       |
      |            +------------------------------+
      |
      +--------------------------------------------> SPEED
         deployment frequency (higher is better)
         lead time for changes (lower is better)

READ THEM IN PAIRS
  frequency alone    -> ship anything, break production
  failure rate alone -> ship nothing, stay perfectly green

LEADING INDICATORS
  pipeline duration and flake rate degrade first

လက်တွေ့ scenario နဲ့ ချိတ်ကြည့်မယ်

DORA ကို တိုင်းဖို့ တတိယအဖွဲ့ tool မလိုပါဘူး — လိုတာက တစ်သမတ်တည်း မှတ်တမ်းတင်ထားတဲ့ event တွေပါ။ ဒီ workflow က deploy တစ်ခုပြီးတိုင်း event တစ်ခု ထုတ်ပေးတယ်။ record job မှာ if: always() ရှိတာ အရေးအကြီးဆုံးပါ — deploy ကျတဲ့အခါလည်း မှတ်တမ်းတင်ရမယ်၊ မဟုတ်ရင် အောင်မြင်တဲ့ deploy တွေကိုပဲ ရေတွက်နေပြီး change failure rate က ထာဝရ သုည ဖြစ်နေမယ်။ ဒါက လက်တွေ့မှာ အဖြစ်အများဆုံး တိုင်းတာမှု အမှားပါ။

needs.deploy.result ကို ရယူပြီး status အဖြစ် ပို့ထားတယ်၊ ဒါကြောင့် event တစ်ခုချင်းစီမှာ အောင်ခဲ့လား ကျခဲ့လား ပါလာမယ် — အဲဒီ နှစ်ခုကနေ deployment frequency နဲ့ change failure rate နှစ်ခုလုံးကို တွက်လို့ရတယ်။ fetch-depth: 0 က commit history အပြည့် ရစေလို့ lead time ကို တွက်တဲ့အခါ ဒီ deploy ထဲ ပါလာတဲ့ commit တွေရဲ့ author date ကို ဖတ်နိုင်တယ် — အဲဒါ မရှိရင် shallow clone ကြောင့် commit တစ်ခုတည်းပဲ မြင်ရပြီး lead time က မှားပါလိမ့်မယ်။

အရေးကြီးတာ တစ်ခု ကျန်သေးတယ် — time to restore ကို ဒီ workflow တစ်ခုတည်းက မတိုင်းနိုင်ဘူး။ အဲဒါက incident စတဲ့အချိန်ကို လိုအပ်ပြီး၊ အဲဒီအချိန်က alerting system ဒါမှမဟုတ် incident tracker ထဲမှာ ရှိတယ်။ ဒါကြောင့် metric လေးခုလုံး လိုချင်ရင် deploy event တွေကို incident record တွေနဲ့ ချိတ်ဆက်ရပါလိမ့်မယ်။

အတူတူ စမ်းရေးကြည့်မယ်

yaml
name: Deploy and Record DORA Events
on:
  push:
    branches:
      - main
permissions:
  contents: read
jobs:
  deploy:
    runs-on: ubuntu-latest
    environment:
      name: production
      url: https://example.com
    steps:
      - uses: actions/checkout@v4
      - name: Deploy the current commit to production
        run: ./scripts/deploy.sh production
      - name: Verify the deployment is serving traffic
        run: ./scripts/smoke.sh https://example.com
  record:
    needs: deploy
    if: always()
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - name: Record the deployment event including failures
        run: ./scripts/record-deploy-event.sh --sha ${{ github.sha }} --status ${{ needs.deploy.result }}
        env:
          METRICS_TOKEN: ${{ secrets.METRICS_TOKEN }}
      - name: Compute lead time from commit dates in this deploy
        run: ./scripts/compute-lead-time.sh --sha ${{ github.sha }}
      - uses: actions/upload-artifact@v4
        with:
          name: dora-events
          path: metrics/
          retention-days: 90
You should see
main ဆီ push တိုင်းမှာ deploy job က production environment အောက်မှာ run ပြီး smoke test နဲ့ အတည်ပြုတယ်။ ပြီးရင် record job က if: always() ကြောင့် deploy အောင်ရင်ရော ကျရင်ရော run တယ်၊ ပြီးတော့ needs.deploy.result ကနေ ရလာတဲ့ status ကို event ထဲ ထည့်ပို့တယ်။ ဒါကြောင့် metric store ထဲမှာ ကြိုးစားခဲ့တဲ့ deployment အားလုံး — အောင်တာရော ကျတာရော — မှတ်တမ်း ရှိလာပြီး deployment frequency နဲ့ change failure rate နှစ်ခုလုံးကို တွက်လို့ ရလာမယ်။ fetch-depth: 0 ကြောင့် commit history အပြည့်ရလို့ lead time တွက်ချက်မှုက ဒီ deploy ထဲပါတဲ့ commit တွေအားလုံးကို မြင်ရတယ်။ deploy job ကို production environment နဲ့ ချိတ်ထားလို့ အဲဒီမှာ approval ထားရင် စောင့်ချိန်ဟာ lead time ထဲ ကျရောက်ပါလိမ့်မယ်။ time to restore ကတော့ ဒီ workflow ထဲမှာ မပါဘဲ incident record တွေနဲ့ ချိတ်မှသာ ရပါမယ်။

၅ မိနစ် စမ်းကြည့်

လွန်ခဲ့တဲ့ ရက် ၉၀ အတွင်း ကိုယ့် repository ရဲ့ deploy workflow run တွေကို API ကနေ ဆွဲထုတ်ပြီး deployment frequency နဲ့ change failure rate ကို တွက်ကြည့်ပါ (ကျသွားတဲ့ run တွေကို မမေ့ပါနဲ့)။ ပြီးရင် အဲဒီ ကာလအတွင်း incident တွေရဲ့ စချိန်နဲ့ ပြန်ကောင်းချိန်ကို လက်နဲ့ ဖြစ်ဖြစ် ထည့်ပြီး time to restore ကို တွက်ပါ။ နောက်ဆုံး "deployment တစ်ခု" နဲ့ "failure တစ်ခု" ဆိုတာ ဘာကို ဆိုလိုတယ်ဆိုတဲ့ အဓိပ္ပာယ်ကို စာတစ်မျက်နှာအောက်နဲ့ ရေးမှတ်ပြီး အသင်းနဲ့ သဘောတူထားပါ။

သတိလေးတစ်ချက်

အောင်မြင်တဲ့ deploy တွေကိုပဲ မှတ်တမ်းတင်ထားလို့ change failure rate က ထာဝရ သုည ဖြစ်နေတာ — dashboard က လှနေပေမယ့် ဘာမှ မပြောပြဘူး

DORA metric တွေကို developer တစ်ဦးချင်းရဲ့ စွမ်းဆောင်ရည် ပစ်မှတ် အဖြစ် သတ်မှတ်လိုက်တာ — PR တွေကို ပိုသေးအောင် ခွဲပြီး deploy အရေအတွက် ဖောင်းပွစေတာမျိုးနဲ့ ကိန်းဂဏန်းက အဓိပ္ပာယ် ကုန်သွားတယ်

DORA: The four keys of software delivery performanceCI/CD with GitHub Actions

ဒီနေရာမှာ လူအများမှားတတ်တယ်

  • အောင်မြင်တဲ့ deploy တွေကိုပဲ မှတ်တမ်းတင်ထားလို့ change failure rate က ထာဝရ သုည ဖြစ်နေတာ — dashboard က လှနေပေမယ့် ဘာမှ မပြောပြဘူး
  • DORA metric တွေကို developer တစ်ဦးချင်းရဲ့ စွမ်းဆောင်ရည် ပစ်မှတ် အဖြစ် သတ်မှတ်လိုက်တာ — PR တွေကို ပိုသေးအောင် ခွဲပြီး deploy အရေအတွက် ဖောင်းပွစေတာမျိုးနဲ့ ကိန်းဂဏန်းက အဓိပ္ပာယ် ကုန်သွားတယ်
  • Workflow ကို production branch ပေါ် တိုက်ရိုက်မစမ်းဘဲ branch (သို့) test repository တွင် အရင်အတည်ပြုပါ။

လေ့ကျင့်ခန်း

လွန်ခဲ့တဲ့ ရက် ၉၀ အတွင်း ကိုယ့် repository ရဲ့ deploy workflow run တွေကို API ကနေ ဆွဲထုတ်ပြီး deployment frequency နဲ့ change failure rate ကို တွက်ကြည့်ပါ (ကျသွားတဲ့ run တွေကို မမေ့ပါနဲ့)။ ပြီးရင် အဲဒီ ကာလအတွင်း incident တွေရဲ့ စချိန်နဲ့ ပြန်ကောင်းချိန်ကို လက်နဲ့ ဖြစ်ဖြစ် ထည့်ပြီး time to restore ကို တွက်ပါ။ နောက်ဆုံး "deployment တစ်ခု" နဲ့ "failure တစ်ခု" ဆိုတာ ဘာကို ဆိုလိုတယ်ဆိုတဲ့ အဓိပ္ပာယ်ကို စာတစ်မျက်နှာအောက်နဲ့ ရေးမှတ်ပြီး အသင်းနဲ့ သဘောတူထားပါ။

You'll know it worked when: main ဆီ push တိုင်းမှာ deploy job က production environment အောက်မှာ run ပြီး smoke test နဲ့ အတည်ပြုတယ်။ ပြီးရင် record job က if: always() ကြောင့် deploy အောင်ရင်ရော ကျရင်ရော run တယ်၊ ပြီးတော့ needs.deploy.result ကနေ ရလာတဲ့ status ကို event ထဲ ထည့်ပို့တယ်။ ဒါကြောင့် metric store ထဲမှာ ကြိုးစားခဲ့တဲ့ deployment အားလုံး — အောင်တာရော ကျတာရော — မှတ်တမ်း ရှိလာပြီး deployment frequency နဲ့ change failure rate နှစ်ခုလုံးကို တွက်လို့ ရလာမယ်။ fetch-depth: 0 ကြောင့် commit history အပြည့်ရလို့ lead time တွက်ချက်မှုက ဒီ deploy ထဲပါတဲ့ commit တွေအားလုံးကို မြင်ရတယ်။ deploy job ကို production environment နဲ့ ချိတ်ထားလို့ အဲဒီမှာ approval ထားရင် စောင့်ချိန်ဟာ lead time ထဲ ကျရောက်ပါလိမ့်မယ်။ time to restore ကတော့ ဒီ workflow ထဲမှာ မပါဘဲ incident record တွေနဲ့ ချိတ်မှသာ ရပါမယ်။

DORA နဲ့ Pipeline ကျန်းမာရေးကို တိုင်းတာခြင်း | Thuta Learning