Lab 19 — GitOps với Monitoring: ArgoCD + Prometheus + Grafana
🎯 1. Mục tiêu
Sau Lab này, bạn sẽ hiểu và thực hành được:
- Vì sao GitOps không chỉ cần deploy, mà còn cần monitoring.
- Cách kết hợp ArgoCD + Prometheus + Grafana.
- Theo dõi trạng thái Application được ArgoCD quản lý.
- Xây dựng dashboard monitoring cơ bản cho GitOps.
- Phân biệt Application Health và Application Metrics.
- Hiểu cách thiết kế monitoring GitOps trong môi trường production.
🤔 2. Vấn đề thực tế
Ở các Lab trước, chúng ta đã có flow:
Developer
│
│ git push
▼
GitOps Repository
│
▼
ArgoCD
│
│ Sync
▼
Kubernetes
ArgoCD giúp chúng ta biết:
"Kubernetes hiện tại có giống với trạng thái mong muốn trong Git không?"
Nhưng production còn một câu hỏi khác:
"Application sau khi deploy có thực sự đang hoạt động tốt không?"
Ví dụ:
Git
│
▼
ArgoCD
│
▼
Kubernetes
│
▼
Todo API
ArgoCD có thể báo:
Application: Synced
Health: Healthy
Nhưng thực tế:
HTTP 500 ↑
CPU ↑
Memory ↑
Request latency ↑
Pod restart ↑
Application vẫn có thể được ArgoCD đánh dấu Healthy, nhưng người dùng lại không truy cập được.
Đây là lý do chúng ta cần kết hợp:
ArgoCD
│
│ Deployment State
▼
Kubernetes
Prometheus
│
│ Runtime Metrics
▼
Grafana
│
▼
DevOps Engineer
ArgoCD cho biết hệ thống đã deploy đúng chưa. Prometheus cho biết hệ thống đang chạy tốt không.
🧠 3. Hiểu nhanh
3.1. GitOps Monitoring là gì?
GitOps Monitoring là việc theo dõi cả:
- Deployment State
- Application Health
- Infrastructure Metrics
- GitOps/ArgoCD Metrics
Có thể hình dung:
Git
│
▼
┌───────┐
│ ArgoCD│
└───┬───┘
│
"Deploy đúng chưa?"
│
▼
Kubernetes
│
"App chạy tốt không?"
│
┌─────────┴─────────┐
▼ ▼
Prometheus ArgoCD
│ Metrics
│
▼
Grafana
│
▼
DevOps Engineer
3.2. ArgoCD và Prometheus khác nhau thế nào?
Đây là điểm người mới rất dễ nhầm.
ArgoCD
ArgoCD quan tâm:
Git State
↕
Cluster State
Ví dụ:
Git:
replicas: 3
Kubernetes:
replicas: 3
→ Synced
Prometheus
Prometheus quan tâm:
Application Runtime
Ví dụ:
CPU: 85%
Memory: 72%
HTTP 500: 120/min
Latency: 2.3s
Pod Restart: 5
Vì vậy:
| Công cụ | Câu hỏi |
|---|---|
| Git | Muốn hệ thống như thế nào? |
| ArgoCD | Cluster có giống Git không? |
| Prometheus | Hệ thống đang hoạt động thế nào? |
| Grafana | Làm sao nhìn các metrics dễ dàng? |
🏗️ 4. Architecture
Trong Lab này, chúng ta sử dụng architecture:
┌──────────────┐
│ GitOps Repo │
└──────┬───────┘
│
│ Watch
▼
┌──────────────┐
│ ArgoCD │
└──────┬───────┘
│
│ Sync
▼
┌──────────────────────┐
│ Kubernetes │
│ │
│ ┌───────────────┐ │
│ │ Todo App │ │
│ └───────────────┘ │
│ │
│ ┌───────────────┐ │
│ │ Prometheus │ │
│ └───────┬───────┘ │
│ │ │
│ │ Metrics │
│ ▼ │
│ ┌───────────────┐ │
│ │ Grafana │ │
│ └───────────────┘ │
└──────────────────────┘
Flow chính:
Git
↓
ArgoCD
↓
Kubernetes
↓
Application
↓
Metrics
↓
Prometheus
↓
Grafana
🛠️ 5. Chuẩn bị môi trường
Lab này giả sử bạn đã có:
- Kubernetes cluster
- ArgoCD
- Todo Application
- Prometheus
- Grafana
Kiểm tra Kubernetes:
kubectl get nodes
Kiểm tra ArgoCD:
kubectl get pods -n argocd
Kiểm tra monitoring:
kubectl get pods -n monitoring
Bạn nên thấy các Pod tương tự:
argocd-server
argocd-repo-server
argocd-application-controller
...
prometheus-server
grafana
...
🔎 6. Kiểm tra ArgoCD Application
💡 Explain
Trước tiên chúng ta cần biết ArgoCD đang quản lý Application nào.
Chạy:
kubectl get applications -n argocd
Ví dụ:
NAME SYNC STATUS HEALTH STATUS
todo-app Synced Healthy
Ở đây:
Synced
Có nghĩa:
Git State
=
Cluster State
Healthy
Có nghĩa ArgoCD đánh giá resource đang ở trạng thái hoạt động bình thường.
🧪 Do
Kiểm tra chi tiết:
kubectl get application todo-app -n argocd
Hoặc:
kubectl describe application todo-app -n argocd
👀 See Result
Bạn có thể thấy:
Sync Status: Synced
Health Status: Healthy
🧠 Understand
Đây mới chỉ là deployment monitoring.
Ví dụ:
ArgoCD:
Synced
Healthy
không đồng nghĩa với:
CPU OK
Memory OK
Latency OK
HTTP 500 OK
Để biết những điều này chúng ta cần Prometheus.
📊 7. Kiểm tra Prometheus
💡 Explain
Prometheus là hệ thống dùng để thu thập và lưu trữ metrics.
Ví dụ application expose:
http_requests_total 1000
http_request_duration_seconds 0.2
Prometheus định kỳ lấy các metrics này và lưu lại.
Flow:
Application
│
│ /metrics
▼
Prometheus
│
│ Query
▼
Grafana
🧪 Do
Kiểm tra Prometheus:
kubectl get pods -n monitoring
Sau đó port-forward:
kubectl port-forward svc/prometheus-server 9090:80 -n monitoring
Mở:
http://localhost:9090
🔍 8. Kiểm tra Kubernetes Metrics
Trong Prometheus thử query:
up
Query này giúp kiểm tra target nào đang hoạt động.
Bạn có thể thử:
container_cpu_usage_seconds_total
hoặc:
container_memory_working_set_bytes
👀 See Result
Nếu Prometheus đang scrape Kubernetes metrics, bạn sẽ thấy nhiều time series.
Ví dụ:
container_cpu_usage_seconds_total
├── pod=todo-backend
├── pod=todo-frontend
└── pod=postgres
🧠 Understand
Điều quan trọng không phải là nhớ tên từng metric.
Bạn chỉ cần nhớ:
Metric
│
▼
Prometheus
│
▼
PromQL
PromQL là ngôn ngữ dùng để truy vấn metrics trong Prometheus.
📈 9. Kết nối Grafana với Prometheus
💡 Explain
Prometheus rất mạnh trong việc lưu trữ và query metrics.
Nhưng giao diện của Prometheus không phải công cụ dashboard tốt nhất.
Grafana giải quyết vấn đề đó.
Prometheus
│
│ Metrics
▼
Grafana
│
▼
Dashboard
🧪 Do
Port-forward Grafana:
kubectl port-forward svc/grafana 3000:80 -n monitoring
Mở:
http://localhost:3000
Đăng nhập bằng tài khoản Grafana của lab trước.
Kiểm tra Data Source:
Connections
↓
Data Sources
↓
Prometheus
URL thường có dạng:
http://prometheus-server.monitoring.svc.cluster.local
📊 10. Tạo GitOps Monitoring Dashboard
💡 Explain
Thay vì mỗi lần cần kiểm tra phải chạy:
kubectl get pods
kubectl get applications
kubectl top pods
chúng ta tạo một dashboard.
Dashboard nên trả lời nhanh:
Deployment có ổn không?
Application có lỗi không?
Pod có restart không?
CPU/Memory có tăng không?
ArgoCD có Sync lỗi không?
🧪 Do
Trong Grafana:
Dashboards
↓
New Dashboard
↓
Add visualization
Tạo một số panel cơ bản.
Panel 1 — Pod CPU
sum by (pod) (
rate(container_cpu_usage_seconds_total[5m])
)
Panel 2 — Pod Memory
sum by (pod) (
container_memory_working_set_bytes
)
Panel 3 — Pod Restart
sum by (pod) (
kube_pod_container_status_restarts_total
)
Panel 4 — Kubernetes Pods
count(kube_pod_info)
🚀 11. Monitoring ArgoCD
💡 Explain
Đây là phần quan trọng nhất của Lab.
Chúng ta không chỉ monitoring Application.
Chúng ta còn muốn biết:
ArgoCD có đang hoạt động bình thường không?
Ví dụ:
Application Sync Failed
Application OutOfSync
ArgoCD API errors
Sync duration tăng
Nếu ArgoCD gặp vấn đề, GitOps pipeline có thể bị ảnh hưởng.
🧠 ArgoCD Metrics
ArgoCD cung cấp metrics để monitoring chính ArgoCD.
Một số nhóm metrics liên quan đến:
Application
Sync
Health
Controller
API Server
Repo Server
Architecture:
ArgoCD
┌───────────┐
│ Controller│
│ API Server│
│ RepoServer│
└─────┬─────┘
│
Metrics
│
▼
Prometheus
│
▼
Grafana
🔎 12. Kiểm tra ArgoCD Metrics
🧪 Do
Kiểm tra Service:
kubectl get svc -n argocd
Bạn sẽ thấy các service liên quan đến ArgoCD.
Sau đó kiểm tra metrics endpoint của component phù hợp trong cluster.
Ví dụ:
kubectl get svc -n argocd
và:
kubectl get pods -n argocd
👀 See Result
Mục tiêu là xác định được:
ArgoCD Component
│
│ /metrics
▼
Prometheus
Nếu Prometheus đã được cấu hình scrape ArgoCD, bạn có thể query các metrics ArgoCD trong Prometheus.
📉 13. Monitoring GitOps Drift
💡 Explain
Một trong những vấn đề GitOps quan trọng nhất là drift.
Drift nghĩa là:
Kubernetes thực tế đã khác với trạng thái được định nghĩa trong Git.
Ví dụ Git:
replicas: 3
Nhưng ai đó chạy:
kubectl scale deployment todo-backend --replicas=5
Cluster:
replicas: 5
Git:
replicas: 3
→ Drift.
🧪 Do
Thử thay đổi trực tiếp:
kubectl scale deployment todo-backend \
--replicas=5 \
-n todo-app
Sau đó kiểm tra ArgoCD:
kubectl get application todo-app -n argocd
👀 See Result
ArgoCD có thể phát hiện:
Sync Status: OutOfSync
Đây là một trong những giá trị lớn nhất của GitOps.
Git
│
│ Desired State
▼
ArgoCD
│
│ Compare
▼
Kubernetes
│
│ Actual State
▼
Drift detected
🔄 14. Self-Healing
💡 Explain
Nếu bật:
automated sync
+
selfHeal
ArgoCD có thể tự đưa cluster trở về trạng thái trong Git.
Ví dụ:
Git
replicas = 3
│
▼
ArgoCD
│
▼
Kubernetes
replicas = 5
│
▼
Drift detected
│
▼
ArgoCD
│
▼
replicas = 3
Điều này rất hữu ích khi Kubernetes bị thay đổi ngoài GitOps workflow.
🧪 Do
Khôi phục deployment về trạng thái GitOps:
kubectl scale deployment todo-backend \
--replicas=5 \
-n todo-app
Sau đó quan sát:
kubectl get application todo-app -n argocd
Nếu Self-Healing được bật, ArgoCD sẽ phát hiện và reconcile lại.
Kiểm tra:
kubectl get deployment todo-backend -n todo-app
👀 See Result
Cuối cùng:
Desired State = Git
Actual State = Kubernetes
và:
Synced
Healthy
🚨 15. Tạo Alert cho GitOps
💡 Explain
Dashboard giúp chúng ta nhìn thấy vấn đề.
Alert giúp chúng ta biết khi vấn đề xảy ra mà không cần ngồi nhìn dashboard.
Ví dụ production:
02:00 AM
ArgoCD Application
│
▼
OutOfSync
│
▼
Alert
│
▼
Slack / Email / PagerDuty
│
▼
DevOps Engineer
🎯 Những alert nên có
Trong production, có thể bắt đầu với:
Application OutOfSync
Application Health Degraded
Application Sync Failed
ArgoCD component unavailable
Prometheus target down
High CPU
High Memory
High Error Rate
Không nên tạo hàng trăm alert ngay từ đầu.
Alert tốt là alert dẫn tới một hành động cụ thể.
Nếu alert chỉ khiến team nhận hàng trăm notification nhưng không biết phải làm gì → alert đó chưa tốt.
🧪 16. Exercise — Tạo GitOps Monitoring Dashboard
Bây giờ hãy tự xây dashboard với ít nhất 6 panels.
Panel 1
ArgoCD Application Status
Panel 2
ArgoCD Application Health
Panel 3
Pod CPU
Panel 4
Pod Memory
Panel 5
Pod Restart
Panel 6
HTTP Error Rate
Dashboard cuối cùng có thể trông như:
┌───────────────────────────────────────────────┐
│ GitOps Monitoring Dashboard │
├─────────────────┬─────────────────────────────┤
│ ArgoCD Sync │ Application Health │
│ Healthy │ Healthy │
├─────────────────┼─────────────────────────────┤
│ CPU │ Memory │
│ 45% │ 62% │
├─────────────────┼─────────────────────────────┤
│ Pod Restart │ HTTP Error Rate │
│ 0 │ 0.2% │
└─────────────────┴─────────────────────────────┘
🧪 17. Exercise — Tạo tình huống lỗi
Đây là phần quan trọng nhất.
Không chỉ tạo dashboard rồi nhìn nó.
Hãy tạo lỗi và xem monitoring phản ứng như thế nào.
Bước 1 — Scale application
kubectl scale deployment todo-backend \
--replicas=5 \
-n todo-app
Bước 2 — Kiểm tra ArgoCD
kubectl get application todo-app -n argocd
Bước 3 — Kiểm tra Kubernetes
kubectl get pods -n todo-app
Bước 4 — Kiểm tra Grafana
Quan sát:
CPU
Memory
Pod Count
Pod Restart
Application Status
Bước 5 — Để ArgoCD Self-Healing xử lý
Quan sát quá trình:
Manual Change
↓
Drift
↓
ArgoCD Detect
↓
Sync
↓
Cluster Recovered
🧠 18. Hiểu bản chất sau Lab
Sau Lab này, hãy nhớ mô hình đơn giản:
┌─────────────┐
│ Git │
│ Desired │
│ State │
└──────┬──────┘
│
▼
┌─────────────┐
│ ArgoCD │
│ │
│ Sync/Drift │
└──────┬──────┘
│
▼
┌─────────────┐
│ Kubernetes │
│ Actual │
│ State │
└──────┬──────┘
│
│ Metrics
▼
┌─────────────┐
│ Prometheus │
└──────┬──────┘
│
▼
┌─────────────┐
│ Grafana │
└─────────────┘
Có thể nhớ bằng 3 câu:
Git nói hệ thống phải như thế nào.
ArgoCD đảm bảo Kubernetes đi theo Git.
Prometheus + Grafana cho biết hệ thống đang hoạt động như thế nào.
💡 19. Kinh nghiệm Production
19.1. Đừng chỉ monitoring Kubernetes
Một hệ thống:
Pod = Running
không có nghĩa:
Application = Healthy
Bạn nên monitoring cả:
Infrastructure
+
Kubernetes
+
Application
+
GitOps
19.2. Đừng biến Grafana thành "bức tường biểu đồ"
Dashboard production không cần 50 biểu đồ.
Hãy ưu tiên:
Golden Signals
├── Latency
├── Traffic
├── Errors
└── Saturation
và GitOps signals:
Sync Status
Health Status
Sync Failures
Drift
19.3. Alert phải có hành động
Ví dụ tốt:
Todo API error rate > 5%
for 5 minutes
→ Có thể điều tra.
Ví dụ không tốt:
CPU = 61%
→ CPU 61% chưa chắc là vấn đề.
19.4. GitOps không thay thế Monitoring
Đây là một misconception khá phổ biến.
GitOps
≠
Monitoring
GitOps giải quyết:
Deployment
Configuration
Drift
Reconciliation
Monitoring giải quyết:
Health
Performance
Errors
Capacity
Availability
Hai hệ thống bổ sung cho nhau.
🏁 20. Tổng kết
Trong Lab này, chúng ta đã xây dựng:
GitOps
│
▼
ArgoCD
│
├── Sync Status
├── Health Status
└── Drift Detection
│
▼
Kubernetes
│
▼
Prometheus
│
▼
Grafana
Quan trọng nhất, bạn đã chuyển từ tư duy:
"Deploy thành công"
sang:
"Deploy đúng"
+
"Application đang khỏe"
+
"Biết ngay khi có vấn đề"
Đó mới là tư duy cần có khi xây dựng một hệ thống Production GitOps.
🎯 21. Checklist
- [ ] Hiểu sự khác nhau giữa ArgoCD và Prometheus
- [ ] Kiểm tra được ArgoCD Application
- [ ] Kiểm tra được Prometheus
- [ ] Kết nối Grafana với Prometheus
- [ ] Tạo GitOps Monitoring Dashboard
- [ ] Theo dõi CPU / Memory / Pod Restart
- [ ] Theo dõi Application Sync / Health
- [ ] Hiểu GitOps Drift
- [ ] Hiểu Self-Healing
- [ ] Biết những loại Alert quan trọng trong production
Lab tiếp theo: chúng ta sẽ đi sâu hơn vào ArgoCD Security & RBAC — thay vì để mọi người có toàn quyền với toàn bộ Kubernetes cluster, chúng ta sẽ học cách kiểm soát ai được phép deploy gì, vào đâu và bằng quyền nào.
All rights reserved