Lab 8.4 - Thiết lập Alerting với Prometheus và Alertmanager
Phần 6 - Kiểm thử Alert trong thực tế (Pending → Firing → Resolved)
Ở các phần trước, chúng ta đã hoàn thành việc xây dựng hệ thống Alerting:
- Prometheus thu thập Metrics.
- Alert Rule được cấu hình.
- Prometheus đánh giá điều kiện Alert.
- Alertmanager nhận Alert.
- Alertmanager có thể gửi Email/Slack.
Tuy nhiên, một hệ thống Alerting chỉ thực sự có giá trị khi chúng ta biết nó hoạt động trong tình huống thực tế.
Trong phần này, chúng ta sẽ tạo ra các sự cố giả lập trên hệ thống Todo và quan sát toàn bộ vòng đời của một Alert:
Inactive
↓
Pending
↓
Firing
↓
Resolved
Mục tiêu
Sau phần này, bạn sẽ hiểu:
- Cách kiểm tra Alert Rule có hoạt động hay không.
- Sự khác nhau giữa Pending và Firing.
- Cách Prometheus gửi Alert sang Alertmanager.
- Cách xác nhận Email/Slack Notification.
- Cách xử lý khi Alert không kích hoạt.
Chuẩn bị môi trường
Trước khi test, kiểm tra toàn bộ hệ thống:
Kubernetes
kubectl get pods -n todo-app
Ví dụ:
NAME READY STATUS
todo-backend 1/1 Running
todo-frontend 1/1 Running
todo-postgres 1/1 Running
Monitoring
kubectl get pods -n monitoring
Cần đảm bảo:
grafana Running
prometheus-server Running
prometheus-alertmanager Running
prometheus-kube-state-metrics Running
Kiểm tra Alert hiện tại
Truy cập Prometheus:
kubectl port-forward \
svc/prometheus-server \
9090:80 \
-n monitoring
Mở:
http://localhost:9090
Vào:
Alerts
Hiện tại:
TodoApiDown
State: Inactive
Điều này đúng.
Vì:
up{job="todo-api"}
đang trả về:
1
Có nghĩa:
Prometheus
↓
todo-backend-service
↓
/actuator/prometheus
OK
Test 1 - Backend Down Alert
Đây là Alert quan trọng nhất trong Production.
Tình huống:
Backend bị crash hoặc không thể phục vụ request.
Bước 1: Scale Backend về 0
Chạy:
kubectl scale deployment todo-backend \
--replicas=0 \
-n todo-app
Kiểm tra:
kubectl get pods -n todo-app
Kết quả:
todo-backend
Terminating
hoặc không còn Pod Backend.
Bước 2: Kiểm tra Metric up
Trong Prometheus:
Query:
up{job="todo-api"}
Trước khi lỗi:
1
Sau khi Backend mất:
0
Điều này có nghĩa:
Prometheus không scrape được Application Metrics.
Alert chuyển sang Pending
Trong Alert Rule:
for: 1m
nghĩa là:
Điều kiện phải đúng liên tục 1 phút.
Ngay sau khi Backend down:
up == 0
Alert chưa gửi ngay.
Trạng thái:
TodoApiDown
PENDING
Vì sao cần Pending?
Nếu không có for:
Ví dụ:
Pod restart 5 giây
Prometheus có thể gửi cảnh báo ngay.
Điều này tạo ra rất nhiều Alert giả.
for giúp xác nhận:
"Đây thực sự là một sự cố."
Alert chuyển sang Firing
Sau 1 phút:
Điều kiện vẫn đúng:
up{job="todo-api"} == 0
Alert chuyển:
TodoApiDown
FIRING
Luồng lúc này:
Prometheus
|
|
▼
Alert Rule
|
|
▼
Alertmanager
|
|
▼
Email / Slack
Kiểm tra Alertmanager
Port Forward:
kubectl port-forward \
svc/prometheus-alertmanager \
9093:9093 \
-n monitoring
Mở:
http://localhost:9093
Vào:
Alerts
Bạn sẽ thấy:
TodoApiDown
Status:
Firing
Thông tin bao gồm:
- Alert Name.
- Severity.
- Namespace.
- Instance.
- Start Time.
Khôi phục Backend
Scale lại:
kubectl scale deployment todo-backend \
--replicas=1 \
-n todo-app
Kiểm tra:
kubectl get pods -n todo-app
Kết quả:
todo-backend
Running
Alert chuyển sang Resolved
Sau khi Prometheus scrape lại thành công:
Metric:
up{job="todo-api"}
trở về:
1
Alert:
Firing
↓
Resolved
Nếu cấu hình:
send_resolved: true
DevOps nhận được:
[RESOLVED]
Todo API recovered
Test 2 - HTTP 500 Alert
Backend có thể vẫn Running nhưng API bị lỗi.
Đây là tình huống rất phổ biến.
Ví dụ:
- Database lỗi.
- Code Bug.
- External Service lỗi.
Pod vẫn:
Running
nhưng người dùng không sử dụng được.
Tạo HTTP 500
Sửa Controller:
Ví dụ:
@GetMapping("/api/todos")
public List<Todo> getTodos(){
throw new RuntimeException();
}
Build và Deploy lại.
Sau đó gọi:
curl http://localhost:8080/api/todos
Kết quả:
500 Internal Server Error
Kiểm tra Metric
Prometheus Query:
http_server_requests_seconds_count{
status="500"
}
Ví dụ:
25
Kiểm tra Alert
Nếu có Rule:
HighHttp5xxRate
Trạng thái:
Ban đầu:
Inactive
Sau khi đủ điều kiện:
Pending
Sau:
for: 2m
Chuyển:
Firing
Test 3 - Pod Restart Alert
Tình huống:
Ứng dụng bị crash liên tục.
Xóa Pod Backend
Lấy tên Pod:
kubectl get pods -n todo-app
Ví dụ:
todo-backend-7d88f7c9d9-x5abc
Xóa:
kubectl delete pod \
todo-backend-7d88f7c9d9-x5abc \
-n todo-app
Deployment sẽ tự tạo Pod mới.
Kiểm tra Restart Count
kubectl get pods -n todo-app
Ví dụ:
NAME
RESTARTS
todo-backend
5
Prometheus Metric
Query:
kube_pod_container_status_restarts_total
Ví dụ:
5
Alert:
PodRestartTooManyTimes
Firing
Test 4 - CPU cao
Đây là Alert phổ biến nhất.
Tạo tải
Cài Apache Benchmark nếu chưa có:
macOS:
brew install httpd
Chạy:
ab \
-n 50000 \
-c 100 \
http://todo.company.local/api/todos
Giải thích:
-n 50000
Tổng số request
-c 100
100 request đồng thời
Quan sát Dashboard
Mở Grafana.
Theo dõi:
Request Rate
10 req/s
↓
100 req/s
Latency
50ms
↓
300ms
CPU
30%
↓
90%
Nếu Alert:
HighCPUUsage
được cấu hình:
Trạng thái:
Pending
sau đó:
Firing
Debug khi Alert không chạy
Trong thực tế, Alert không hoạt động là lỗi rất thường gặp.
Trường hợp 1: Không thấy Alert trong Prometheus
Kiểm tra:
Status
↓
Rules
Nếu không có Rule:
Kiểm tra:
helm get values prometheus \
-n monitoring
Trường hợp 2: Alert luôn Inactive
Kiểm tra PromQL.
Ví dụ:
up{job="todo-api"}
Nếu:
1
thì hệ thống đang khỏe.
Trường hợp 3: Prometheus không gửi Alertmanager
Kiểm tra:
Prometheus:
Status
↓
Runtime & Build Information
Tìm:
Alertmanagers
Phải có:
prometheus-alertmanager:9093
Trường hợp 4: Alertmanager không gửi Email
Kiểm tra log:
kubectl logs \
prometheus-alertmanager-0 \
-n monitoring
Tìm lỗi:
failed to send email
Tổng kết phần 6
Sau phần này, bạn đã thực hành:
✅ Backend Down Alert ✅ HTTP 500 Alert ✅ Pod Restart Alert ✅ CPU Alert ✅ Quan sát Pending ✅ Quan sát Firing ✅ Quan sát Resolved ✅ Kiểm tra Alertmanager Notification
Luồng hoàn chỉnh:
Application
|
|
▼
Metrics
|
|
▼
Prometheus
|
|
▼
Alert Rule
|
|
▼
Alertmanager
|
|
▼
Email / Slack
|
|
▼
DevOps
All Rights Reserved