0

Lab 19 — GitOps với Monitoring: ArgoCD + Prometheus + Grafana

🎯 1. Mục tiêu

Sau Lab này, bạn sẽ hiểu và thực hành được:

  • Vì sao GitOps không chỉ cần deploy, mà còn cần monitoring.
  • Cách kết hợp ArgoCD + Prometheus + Grafana.
  • Theo dõi trạng thái Application được ArgoCD quản lý.
  • Xây dựng dashboard monitoring cơ bản cho GitOps.
  • Phân biệt Application HealthApplication Metrics.
  • Hiểu cách thiết kế monitoring GitOps trong môi trường production.

🤔 2. Vấn đề thực tế

Ở các Lab trước, chúng ta đã có flow:

Developer
    │
    │ git push
    ▼
GitOps Repository
    │
    ▼
ArgoCD
    │
    │ Sync
    ▼
Kubernetes

ArgoCD giúp chúng ta biết:

"Kubernetes hiện tại có giống với trạng thái mong muốn trong Git không?"

Nhưng production còn một câu hỏi khác:

"Application sau khi deploy có thực sự đang hoạt động tốt không?"

Ví dụ:

Git
 │
 ▼
ArgoCD
 │
 ▼
Kubernetes
 │
 ▼
Todo API

ArgoCD có thể báo:

Application: Synced
Health: Healthy

Nhưng thực tế:

HTTP 500        ↑
CPU             ↑
Memory          ↑
Request latency ↑
Pod restart     ↑

Application vẫn có thể được ArgoCD đánh dấu Healthy, nhưng người dùng lại không truy cập được.

Đây là lý do chúng ta cần kết hợp:

ArgoCD
   │
   │ Deployment State
   ▼
Kubernetes

Prometheus
   │
   │ Runtime Metrics
   ▼
Grafana
   │
   ▼
DevOps Engineer

ArgoCD cho biết hệ thống đã deploy đúng chưa. Prometheus cho biết hệ thống đang chạy tốt không.


🧠 3. Hiểu nhanh

3.1. GitOps Monitoring là gì?

GitOps Monitoring là việc theo dõi cả:

  1. Deployment State
  2. Application Health
  3. Infrastructure Metrics
  4. GitOps/ArgoCD Metrics

Có thể hình dung:

                 Git
                  │
                  ▼
              ┌───────┐
              │ ArgoCD│
              └───┬───┘
                  │
          "Deploy đúng chưa?"
                  │
                  ▼
             Kubernetes
                  │
          "App chạy tốt không?"
                  │
        ┌─────────┴─────────┐
        ▼                   ▼
   Prometheus            ArgoCD
        │                 Metrics
        │
        ▼
     Grafana
        │
        ▼
   DevOps Engineer

3.2. ArgoCD và Prometheus khác nhau thế nào?

Đây là điểm người mới rất dễ nhầm.

ArgoCD

ArgoCD quan tâm:

Git State
    ↕
Cluster State

Ví dụ:

Git:
replicas: 3

Kubernetes:
replicas: 3

→ Synced

Prometheus

Prometheus quan tâm:

Application Runtime

Ví dụ:

CPU: 85%
Memory: 72%
HTTP 500: 120/min
Latency: 2.3s
Pod Restart: 5

Vì vậy:

Công cụ Câu hỏi
Git Muốn hệ thống như thế nào?
ArgoCD Cluster có giống Git không?
Prometheus Hệ thống đang hoạt động thế nào?
Grafana Làm sao nhìn các metrics dễ dàng?

🏗️ 4. Architecture

Trong Lab này, chúng ta sử dụng architecture:

                    ┌──────────────┐
                    │ GitOps Repo  │
                    └──────┬───────┘
                           │
                           │ Watch
                           ▼
                    ┌──────────────┐
                    │    ArgoCD    │
                    └──────┬───────┘
                           │
                           │ Sync
                           ▼
                ┌──────────────────────┐
                │      Kubernetes      │
                │                      │
                │  ┌───────────────┐   │
                │  │   Todo App    │   │
                │  └───────────────┘   │
                │                      │
                │  ┌───────────────┐   │
                │  │   Prometheus  │   │
                │  └───────┬───────┘   │
                │          │            │
                │          │ Metrics    │
                │          ▼            │
                │  ┌───────────────┐   │
                │  │    Grafana    │   │
                │  └───────────────┘   │
                └──────────────────────┘

Flow chính:

Git
 ↓
ArgoCD
 ↓
Kubernetes
 ↓
Application
 ↓
Metrics
 ↓
Prometheus
 ↓
Grafana

🛠️ 5. Chuẩn bị môi trường

Lab này giả sử bạn đã có:

  • Kubernetes cluster
  • ArgoCD
  • Todo Application
  • Prometheus
  • Grafana

Kiểm tra Kubernetes:

kubectl get nodes

Kiểm tra ArgoCD:

kubectl get pods -n argocd

Kiểm tra monitoring:

kubectl get pods -n monitoring

Bạn nên thấy các Pod tương tự:

argocd-server
argocd-repo-server
argocd-application-controller
...

prometheus-server
grafana
...

🔎 6. Kiểm tra ArgoCD Application

💡 Explain

Trước tiên chúng ta cần biết ArgoCD đang quản lý Application nào.

Chạy:

kubectl get applications -n argocd

Ví dụ:

NAME        SYNC STATUS   HEALTH STATUS
todo-app    Synced        Healthy

Ở đây:

Synced

Có nghĩa:

Git State
   =
Cluster State

Healthy

Có nghĩa ArgoCD đánh giá resource đang ở trạng thái hoạt động bình thường.


🧪 Do

Kiểm tra chi tiết:

kubectl get application todo-app -n argocd

Hoặc:

kubectl describe application todo-app -n argocd

👀 See Result

Bạn có thể thấy:

Sync Status:   Synced
Health Status: Healthy

🧠 Understand

Đây mới chỉ là deployment monitoring.

Ví dụ:

ArgoCD:

Synced
Healthy

không đồng nghĩa với:

CPU        OK
Memory     OK
Latency    OK
HTTP 500   OK

Để biết những điều này chúng ta cần Prometheus.


📊 7. Kiểm tra Prometheus

💡 Explain

Prometheus là hệ thống dùng để thu thập và lưu trữ metrics.

Ví dụ application expose:

http_requests_total 1000
http_request_duration_seconds 0.2

Prometheus định kỳ lấy các metrics này và lưu lại.

Flow:

Application
     │
     │ /metrics
     ▼
Prometheus
     │
     │ Query
     ▼
Grafana

🧪 Do

Kiểm tra Prometheus:

kubectl get pods -n monitoring

Sau đó port-forward:

kubectl port-forward svc/prometheus-server 9090:80 -n monitoring

Mở:

http://localhost:9090

🔍 8. Kiểm tra Kubernetes Metrics

Trong Prometheus thử query:

up

Query này giúp kiểm tra target nào đang hoạt động.

Bạn có thể thử:

container_cpu_usage_seconds_total

hoặc:

container_memory_working_set_bytes

👀 See Result

Nếu Prometheus đang scrape Kubernetes metrics, bạn sẽ thấy nhiều time series.

Ví dụ:

container_cpu_usage_seconds_total
    ├── pod=todo-backend
    ├── pod=todo-frontend
    └── pod=postgres

🧠 Understand

Điều quan trọng không phải là nhớ tên từng metric.

Bạn chỉ cần nhớ:

Metric
   │
   ▼
Prometheus
   │
   ▼
PromQL

PromQL là ngôn ngữ dùng để truy vấn metrics trong Prometheus.


📈 9. Kết nối Grafana với Prometheus

💡 Explain

Prometheus rất mạnh trong việc lưu trữ và query metrics.

Nhưng giao diện của Prometheus không phải công cụ dashboard tốt nhất.

Grafana giải quyết vấn đề đó.

Prometheus
    │
    │ Metrics
    ▼
Grafana
    │
    ▼
Dashboard

🧪 Do

Port-forward Grafana:

kubectl port-forward svc/grafana 3000:80 -n monitoring

Mở:

http://localhost:3000

Đăng nhập bằng tài khoản Grafana của lab trước.

Kiểm tra Data Source:

Connections
    ↓
Data Sources
    ↓
Prometheus

URL thường có dạng:

http://prometheus-server.monitoring.svc.cluster.local

📊 10. Tạo GitOps Monitoring Dashboard

💡 Explain

Thay vì mỗi lần cần kiểm tra phải chạy:

kubectl get pods
kubectl get applications
kubectl top pods

chúng ta tạo một dashboard.

Dashboard nên trả lời nhanh:

Deployment có ổn không?
Application có lỗi không?
Pod có restart không?
CPU/Memory có tăng không?
ArgoCD có Sync lỗi không?

🧪 Do

Trong Grafana:

Dashboards
   ↓
New Dashboard
   ↓
Add visualization

Tạo một số panel cơ bản.

Panel 1 — Pod CPU

sum by (pod) (
  rate(container_cpu_usage_seconds_total[5m])
)

Panel 2 — Pod Memory

sum by (pod) (
  container_memory_working_set_bytes
)

Panel 3 — Pod Restart

sum by (pod) (
  kube_pod_container_status_restarts_total
)

Panel 4 — Kubernetes Pods

count(kube_pod_info)

🚀 11. Monitoring ArgoCD

💡 Explain

Đây là phần quan trọng nhất của Lab.

Chúng ta không chỉ monitoring Application.

Chúng ta còn muốn biết:

ArgoCD có đang hoạt động bình thường không?

Ví dụ:

Application Sync Failed
Application OutOfSync
ArgoCD API errors
Sync duration tăng

Nếu ArgoCD gặp vấn đề, GitOps pipeline có thể bị ảnh hưởng.


🧠 ArgoCD Metrics

ArgoCD cung cấp metrics để monitoring chính ArgoCD.

Một số nhóm metrics liên quan đến:

Application
Sync
Health
Controller
API Server
Repo Server

Architecture:

                 ArgoCD
              ┌───────────┐
              │ Controller│
              │ API Server│
              │ RepoServer│
              └─────┬─────┘
                    │
                 Metrics
                    │
                    ▼
                Prometheus
                    │
                    ▼
                  Grafana

🔎 12. Kiểm tra ArgoCD Metrics

🧪 Do

Kiểm tra Service:

kubectl get svc -n argocd

Bạn sẽ thấy các service liên quan đến ArgoCD.

Sau đó kiểm tra metrics endpoint của component phù hợp trong cluster.

Ví dụ:

kubectl get svc -n argocd

và:

kubectl get pods -n argocd

👀 See Result

Mục tiêu là xác định được:

ArgoCD Component
       │
       │ /metrics
       ▼
Prometheus

Nếu Prometheus đã được cấu hình scrape ArgoCD, bạn có thể query các metrics ArgoCD trong Prometheus.


📉 13. Monitoring GitOps Drift

💡 Explain

Một trong những vấn đề GitOps quan trọng nhất là drift.

Drift nghĩa là:

Kubernetes thực tế đã khác với trạng thái được định nghĩa trong Git.

Ví dụ Git:

replicas: 3

Nhưng ai đó chạy:

kubectl scale deployment todo-backend --replicas=5

Cluster:

replicas: 5

Git:

replicas: 3

→ Drift.


🧪 Do

Thử thay đổi trực tiếp:

kubectl scale deployment todo-backend \
  --replicas=5 \
  -n todo-app

Sau đó kiểm tra ArgoCD:

kubectl get application todo-app -n argocd

👀 See Result

ArgoCD có thể phát hiện:

Sync Status: OutOfSync

Đây là một trong những giá trị lớn nhất của GitOps.

Git
 │
 │ Desired State
 ▼
ArgoCD
 │
 │ Compare
 ▼
Kubernetes
 │
 │ Actual State
 ▼
Drift detected

🔄 14. Self-Healing

💡 Explain

Nếu bật:

automated sync
+
selfHeal

ArgoCD có thể tự đưa cluster trở về trạng thái trong Git.

Ví dụ:

Git
replicas = 3
     │
     ▼
ArgoCD
     │
     ▼
Kubernetes
replicas = 5
     │
     ▼
Drift detected
     │
     ▼
ArgoCD
     │
     ▼
replicas = 3

Điều này rất hữu ích khi Kubernetes bị thay đổi ngoài GitOps workflow.


🧪 Do

Khôi phục deployment về trạng thái GitOps:

kubectl scale deployment todo-backend \
  --replicas=5 \
  -n todo-app

Sau đó quan sát:

kubectl get application todo-app -n argocd

Nếu Self-Healing được bật, ArgoCD sẽ phát hiện và reconcile lại.

Kiểm tra:

kubectl get deployment todo-backend -n todo-app

👀 See Result

Cuối cùng:

Desired State = Git
Actual State  = Kubernetes

và:

Synced
Healthy

🚨 15. Tạo Alert cho GitOps

💡 Explain

Dashboard giúp chúng ta nhìn thấy vấn đề.

Alert giúp chúng ta biết khi vấn đề xảy ra mà không cần ngồi nhìn dashboard.

Ví dụ production:

02:00 AM

ArgoCD Application
       │
       ▼
OutOfSync
       │
       ▼
Alert
       │
       ▼
Slack / Email / PagerDuty
       │
       ▼
DevOps Engineer

🎯 Những alert nên có

Trong production, có thể bắt đầu với:

Application OutOfSync
Application Health Degraded
Application Sync Failed
ArgoCD component unavailable
Prometheus target down
High CPU
High Memory
High Error Rate

Không nên tạo hàng trăm alert ngay từ đầu.

Alert tốt là alert dẫn tới một hành động cụ thể.

Nếu alert chỉ khiến team nhận hàng trăm notification nhưng không biết phải làm gì → alert đó chưa tốt.


🧪 16. Exercise — Tạo GitOps Monitoring Dashboard

Bây giờ hãy tự xây dashboard với ít nhất 6 panels.

Panel 1

ArgoCD Application Status

Panel 2

ArgoCD Application Health

Panel 3

Pod CPU

Panel 4

Pod Memory

Panel 5

Pod Restart

Panel 6

HTTP Error Rate

Dashboard cuối cùng có thể trông như:

┌───────────────────────────────────────────────┐
│           GitOps Monitoring Dashboard         │
├─────────────────┬─────────────────────────────┤
│ ArgoCD Sync     │ Application Health          │
│     Healthy     │        Healthy              │
├─────────────────┼─────────────────────────────┤
│ CPU             │ Memory                      │
│     45%         │        62%                  │
├─────────────────┼─────────────────────────────┤
│ Pod Restart     │ HTTP Error Rate             │
│       0         │        0.2%                 │
└─────────────────┴─────────────────────────────┘

🧪 17. Exercise — Tạo tình huống lỗi

Đây là phần quan trọng nhất.

Không chỉ tạo dashboard rồi nhìn nó.

Hãy tạo lỗi và xem monitoring phản ứng như thế nào.

Bước 1 — Scale application

kubectl scale deployment todo-backend \
  --replicas=5 \
  -n todo-app

Bước 2 — Kiểm tra ArgoCD

kubectl get application todo-app -n argocd

Bước 3 — Kiểm tra Kubernetes

kubectl get pods -n todo-app

Bước 4 — Kiểm tra Grafana

Quan sát:

CPU
Memory
Pod Count
Pod Restart
Application Status

Bước 5 — Để ArgoCD Self-Healing xử lý

Quan sát quá trình:

Manual Change
      ↓
Drift
      ↓
ArgoCD Detect
      ↓
Sync
      ↓
Cluster Recovered

🧠 18. Hiểu bản chất sau Lab

Sau Lab này, hãy nhớ mô hình đơn giản:

                 ┌─────────────┐
                 │     Git     │
                 │ Desired     │
                 │ State       │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │   ArgoCD    │
                 │             │
                 │ Sync/Drift  │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │ Kubernetes  │
                 │ Actual      │
                 │ State       │
                 └──────┬──────┘
                        │
                        │ Metrics
                        ▼
                 ┌─────────────┐
                 │ Prometheus  │
                 └──────┬──────┘
                        │
                        ▼
                 ┌─────────────┐
                 │   Grafana   │
                 └─────────────┘

Có thể nhớ bằng 3 câu:

Git nói hệ thống phải như thế nào.

ArgoCD đảm bảo Kubernetes đi theo Git.

Prometheus + Grafana cho biết hệ thống đang hoạt động như thế nào.


💡 19. Kinh nghiệm Production

19.1. Đừng chỉ monitoring Kubernetes

Một hệ thống:

Pod = Running

không có nghĩa:

Application = Healthy

Bạn nên monitoring cả:

Infrastructure
+
Kubernetes
+
Application
+
GitOps

19.2. Đừng biến Grafana thành "bức tường biểu đồ"

Dashboard production không cần 50 biểu đồ.

Hãy ưu tiên:

Golden Signals
├── Latency
├── Traffic
├── Errors
└── Saturation

và GitOps signals:

Sync Status
Health Status
Sync Failures
Drift

19.3. Alert phải có hành động

Ví dụ tốt:

Todo API error rate > 5%
for 5 minutes

→ Có thể điều tra.

Ví dụ không tốt:

CPU = 61%

→ CPU 61% chưa chắc là vấn đề.


19.4. GitOps không thay thế Monitoring

Đây là một misconception khá phổ biến.

GitOps
  ≠
Monitoring

GitOps giải quyết:

Deployment
Configuration
Drift
Reconciliation

Monitoring giải quyết:

Health
Performance
Errors
Capacity
Availability

Hai hệ thống bổ sung cho nhau.


🏁 20. Tổng kết

Trong Lab này, chúng ta đã xây dựng:

GitOps
   │
   ▼
ArgoCD
   │
   ├── Sync Status
   ├── Health Status
   └── Drift Detection
           │
           ▼
      Kubernetes
           │
           ▼
      Prometheus
           │
           ▼
        Grafana

Quan trọng nhất, bạn đã chuyển từ tư duy:

"Deploy thành công"

sang:

"Deploy đúng"
       +
"Application đang khỏe"
       +
"Biết ngay khi có vấn đề"

Đó mới là tư duy cần có khi xây dựng một hệ thống Production GitOps.


🎯 21. Checklist

  • [ ] Hiểu sự khác nhau giữa ArgoCD và Prometheus
  • [ ] Kiểm tra được ArgoCD Application
  • [ ] Kiểm tra được Prometheus
  • [ ] Kết nối Grafana với Prometheus
  • [ ] Tạo GitOps Monitoring Dashboard
  • [ ] Theo dõi CPU / Memory / Pod Restart
  • [ ] Theo dõi Application Sync / Health
  • [ ] Hiểu GitOps Drift
  • [ ] Hiểu Self-Healing
  • [ ] Biết những loại Alert quan trọng trong production

Lab tiếp theo: chúng ta sẽ đi sâu hơn vào ArgoCD Security & RBAC — thay vì để mọi người có toàn quyền với toàn bộ Kubernetes cluster, chúng ta sẽ học cách kiểm soát ai được phép deploy gì, vào đâu và bằng quyền nào.


All rights reserved

Viblo
Hãy đăng ký một tài khoản Viblo để nhận được nhiều bài viết thú vị hơn.
Đăng kí