0

Lab 8.4 - Thiết lập Alerting với Prometheus và Alertmanager

Phần 6 - Kiểm thử Alert trong thực tế (Pending → Firing → Resolved)

Ở các phần trước, chúng ta đã hoàn thành việc xây dựng hệ thống Alerting:

  • Prometheus thu thập Metrics.
  • Alert Rule được cấu hình.
  • Prometheus đánh giá điều kiện Alert.
  • Alertmanager nhận Alert.
  • Alertmanager có thể gửi Email/Slack.

Tuy nhiên, một hệ thống Alerting chỉ thực sự có giá trị khi chúng ta biết nó hoạt động trong tình huống thực tế.

Trong phần này, chúng ta sẽ tạo ra các sự cố giả lập trên hệ thống Todo và quan sát toàn bộ vòng đời của một Alert:

Inactive

    ↓

Pending

    ↓

Firing

    ↓

Resolved

Mục tiêu

Sau phần này, bạn sẽ hiểu:

  • Cách kiểm tra Alert Rule có hoạt động hay không.
  • Sự khác nhau giữa Pending và Firing.
  • Cách Prometheus gửi Alert sang Alertmanager.
  • Cách xác nhận Email/Slack Notification.
  • Cách xử lý khi Alert không kích hoạt.

Chuẩn bị môi trường

Trước khi test, kiểm tra toàn bộ hệ thống:

Kubernetes

kubectl get pods -n todo-app

Ví dụ:

NAME                              READY   STATUS
todo-backend                      1/1     Running
todo-frontend                     1/1     Running
todo-postgres                     1/1     Running

Monitoring

kubectl get pods -n monitoring

Cần đảm bảo:

grafana                              Running

prometheus-server                    Running

prometheus-alertmanager              Running

prometheus-kube-state-metrics        Running

Kiểm tra Alert hiện tại

Truy cập Prometheus:

kubectl port-forward \
svc/prometheus-server \
9090:80 \
-n monitoring

Mở:

http://localhost:9090

Vào:

Alerts

Hiện tại:

TodoApiDown

State: Inactive

Điều này đúng.

Vì:

up{job="todo-api"}

đang trả về:

1

Có nghĩa:

Prometheus

   ↓

todo-backend-service

   ↓

/actuator/prometheus

OK

Test 1 - Backend Down Alert

Đây là Alert quan trọng nhất trong Production.

Tình huống:

Backend bị crash hoặc không thể phục vụ request.


Bước 1: Scale Backend về 0

Chạy:

kubectl scale deployment todo-backend \
--replicas=0 \
-n todo-app

Kiểm tra:

kubectl get pods -n todo-app

Kết quả:

todo-backend

Terminating

hoặc không còn Pod Backend.


Bước 2: Kiểm tra Metric up

Trong Prometheus:

Query:

up{job="todo-api"}

Trước khi lỗi:

1

Sau khi Backend mất:

0

Điều này có nghĩa:

Prometheus không scrape được Application Metrics.


Alert chuyển sang Pending

Trong Alert Rule:

for: 1m

nghĩa là:

Điều kiện phải đúng liên tục 1 phút.

Ngay sau khi Backend down:

up == 0

Alert chưa gửi ngay.

Trạng thái:

TodoApiDown

PENDING

Vì sao cần Pending?

Nếu không có for:

Ví dụ:

Pod restart 5 giây

Prometheus có thể gửi cảnh báo ngay.

Điều này tạo ra rất nhiều Alert giả.

for giúp xác nhận:

"Đây thực sự là một sự cố."


Alert chuyển sang Firing

Sau 1 phút:

Điều kiện vẫn đúng:

up{job="todo-api"} == 0

Alert chuyển:

TodoApiDown

FIRING

Luồng lúc này:

Prometheus

    |
    |
    ▼

Alert Rule

    |
    |
    ▼

Alertmanager

    |
    |
    ▼

Email / Slack

Kiểm tra Alertmanager

Port Forward:

kubectl port-forward \
svc/prometheus-alertmanager \
9093:9093 \
-n monitoring

Mở:

http://localhost:9093

Vào:

Alerts

Bạn sẽ thấy:

TodoApiDown

Status:

Firing

Thông tin bao gồm:

  • Alert Name.
  • Severity.
  • Namespace.
  • Instance.
  • Start Time.

Khôi phục Backend

Scale lại:

kubectl scale deployment todo-backend \
--replicas=1 \
-n todo-app

Kiểm tra:

kubectl get pods -n todo-app

Kết quả:

todo-backend

Running

Alert chuyển sang Resolved

Sau khi Prometheus scrape lại thành công:

Metric:

up{job="todo-api"}

trở về:

1

Alert:

Firing

↓

Resolved

Nếu cấu hình:

send_resolved: true

DevOps nhận được:

[RESOLVED]

Todo API recovered

Test 2 - HTTP 500 Alert

Backend có thể vẫn Running nhưng API bị lỗi.

Đây là tình huống rất phổ biến.

Ví dụ:

  • Database lỗi.
  • Code Bug.
  • External Service lỗi.

Pod vẫn:

Running

nhưng người dùng không sử dụng được.


Tạo HTTP 500

Sửa Controller:

Ví dụ:

@GetMapping("/api/todos")
public List<Todo> getTodos(){

    throw new RuntimeException();

}

Build và Deploy lại.

Sau đó gọi:

curl http://localhost:8080/api/todos

Kết quả:

500 Internal Server Error

Kiểm tra Metric

Prometheus Query:

http_server_requests_seconds_count{
status="500"
}

Ví dụ:

25

Kiểm tra Alert

Nếu có Rule:

HighHttp5xxRate

Trạng thái:

Ban đầu:

Inactive

Sau khi đủ điều kiện:

Pending

Sau:

for: 2m

Chuyển:

Firing

Test 3 - Pod Restart Alert

Tình huống:

Ứng dụng bị crash liên tục.


Xóa Pod Backend

Lấy tên Pod:

kubectl get pods -n todo-app

Ví dụ:

todo-backend-7d88f7c9d9-x5abc

Xóa:

kubectl delete pod \
todo-backend-7d88f7c9d9-x5abc \
-n todo-app

Deployment sẽ tự tạo Pod mới.


Kiểm tra Restart Count

kubectl get pods -n todo-app

Ví dụ:

NAME

RESTARTS

todo-backend

5

Prometheus Metric

Query:

kube_pod_container_status_restarts_total

Ví dụ:

5

Alert:

PodRestartTooManyTimes

Firing

Test 4 - CPU cao

Đây là Alert phổ biến nhất.


Tạo tải

Cài Apache Benchmark nếu chưa có:

macOS:

brew install httpd

Chạy:

ab \
-n 50000 \
-c 100 \
http://todo.company.local/api/todos

Giải thích:

-n 50000

Tổng số request


-c 100

100 request đồng thời

Quan sát Dashboard

Mở Grafana.

Theo dõi:

Request Rate

10 req/s

↓

100 req/s

Latency

50ms

↓

300ms

CPU

30%

↓

90%

Nếu Alert:

HighCPUUsage

được cấu hình:

Trạng thái:

Pending

sau đó:

Firing

Debug khi Alert không chạy

Trong thực tế, Alert không hoạt động là lỗi rất thường gặp.


Trường hợp 1: Không thấy Alert trong Prometheus

Kiểm tra:

Status

↓

Rules

Nếu không có Rule:

Kiểm tra:

helm get values prometheus \
-n monitoring

Trường hợp 2: Alert luôn Inactive

Kiểm tra PromQL.

Ví dụ:

up{job="todo-api"}

Nếu:

1

thì hệ thống đang khỏe.


Trường hợp 3: Prometheus không gửi Alertmanager

Kiểm tra:

Prometheus:

Status

↓

Runtime & Build Information

Tìm:

Alertmanagers

Phải có:

prometheus-alertmanager:9093

Trường hợp 4: Alertmanager không gửi Email

Kiểm tra log:

kubectl logs \
prometheus-alertmanager-0 \
-n monitoring

Tìm lỗi:

failed to send email

Tổng kết phần 6

Sau phần này, bạn đã thực hành:

✅ Backend Down Alert ✅ HTTP 500 Alert ✅ Pod Restart Alert ✅ CPU Alert ✅ Quan sát Pending ✅ Quan sát Firing ✅ Quan sát Resolved ✅ Kiểm tra Alertmanager Notification

Luồng hoàn chỉnh:

Application

     |
     |
     ▼

Metrics

     |
     |
     ▼

Prometheus

     |
     |
     ▼

Alert Rule

     |
     |
     ▼

Alertmanager

     |
     |
     ▼

Email / Slack

     |
     |
     ▼

DevOps


All Rights Reserved

Viblo
Let's register a Viblo Account to get more interesting posts.