0

Lab 9 - Xây dựng hệ thống Centralized Logging với Fluent Bit, Loki và Grafana

Series: Kubernetes từ Zero đến Production


Tình huống thực tế

Sau Lab 8, hệ thống Todo đã có Alerting.

Ví dụ:

Alert:

HTTP 500 tăng cao

DevOps nhận được cảnh báo.

Bước tiếp theo là tìm nguyên nhân.

Thông thường, DevOps sẽ làm:

kubectl logs <pod-name>

Ví dụ:

kubectl logs todo-backend-xxxx

Nhưng khi hệ thống lớn hơn:

Production

100 Services

500 Pods

Cách này trở nên không khả thi.

DevOps không thể:

  • SSH từng server
  • Tìm từng Pod
  • Kiểm tra từng file log

Cần một nơi tập trung toàn bộ log của hệ thống.

Đó chính là Centralized Logging.


Mục tiêu

Sau bài lab này, bạn sẽ:

  • Hiểu Centralized Logging
  • Hiểu Kubernetes Logging Architecture
  • Triển khai Fluent Bit
  • Triển khai Loki
  • Kết nối Loki với Grafana
  • Thu thập log từ Kubernetes Pod
  • Tìm kiếm log bằng Grafana
  • Debug lỗi ứng dụng thông qua log

Kiến thức cần chuẩn bị

Đã hoàn thành:

  • ✅ Lab 1 → Lab 8

Đã có:

  • Kubernetes Cluster
  • Prometheus
  • Grafana
  • Alertmanager

Logging Architecture

Kiến trúc hoàn chỉnh:

                  Application
                       |
                    stdout
                       |
                Kubernetes Node
                       |
                  Fluent Bit
                       |
                     Loki
                       |
                    Grafana
                       |
                   Log Search

Kubernetes Logging hoạt động như thế nào?

Trong Kubernetes, ứng dụng thường không ghi log vào file.

Ví dụ Spring Boot:

Console Output
      ↓
   stdout

Kubernetes sẽ tự động lưu stdout:

/var/log/containers

Ví dụ:

todo-backend-xxx.log

Fluent Bit sẽ đọc các log này và gửi tới Loki.


Thành phần trong hệ thống

Fluent Bit

Fluent Bit là Log Collector.

Nhiệm vụ:

  • Đọc log container
  • Parse log
  • Thêm Kubernetes Metadata
  • Forward log

Ví dụ:

Log gốc:

ERROR Database connection failed

Sau Fluent Bit:

{
 "pod":"todo-backend",
 "namespace":"todo-app",
 "container":"backend",
 "message":"ERROR Database connection failed"
}

Loki

Loki là Log Aggregation System.

Khác Elasticsearch:

Elasticsearch:

Index toàn bộ nội dung log

Loki:

Lưu log + Label

Ví dụ:

namespace=todo-app

app=backend

Loki được tối ưu cho Kubernetes.


Grafana

Grafana dùng để:

  • Query Log
  • Filter Log
  • Correlate Metrics + Logs

Ví dụ:

Nhìn thấy:

CPU tăng

Sau đó xem:

Backend Logs

để tìm nguyên nhân.


Bước 1. Tạo Namespace Logging

Tạo namespace:

kubectl create namespace logging

Kiểm tra:

kubectl get namespace

Bước 2. Cài Loki

Thêm Helm Repository:

helm repo add grafana \
https://grafana.github.io/helm-charts

Update:

helm repo update

Cài Loki:

helm install loki \
grafana/loki-stack \
-n logging

Bước 3. Kiểm tra Loki

Kiểm tra Pod:

kubectl get pods -n logging

Ví dụ:

loki-0

Running

Kiểm tra Service:

kubectl get svc -n logging

Ví dụ:

loki

ClusterIP

3100

Loki đang chạy.


Bước 4. Cài Fluent Bit

Fluent Bit chạy dạng DaemonSet.

Điều này có nghĩa:

Node 1

Fluent Bit


Node 2

Fluent Bit

Mỗi Node có một Fluent Bit.

Thêm Helm Repository:

helm repo add fluent \
https://fluent.github.io/helm-charts

Update:

helm repo update

Cài:

helm install fluent-bit \
fluent/fluent-bit \
-n logging

Bước 5. Kiểm tra Fluent Bit

kubectl get pods -n logging

Ví dụ:

fluent-bit-xxxxx

Running

Kiểm tra:

kubectl get daemonset -n logging

Kết quả:

fluent-bit

DESIRED

1

Bước 6. Kiểm tra Log Collection

Xem log Fluent Bit:

kubectl logs \
-l app.kubernetes.io/name=fluent-bit \
-n logging

Bạn sẽ thấy:

tailing /var/log/containers

Fluent Bit đang đọc log container.


Bước 7. Kết nối Loki với Grafana

Mở Grafana:

http://localhost:3000

Vào:

Connections
↓
Data Sources
↓
Add Data Source

Chọn:

Loki

URL:

http://loki.logging.svc.cluster.local:3100

Save:

Save & Test

Kết quả:

Data source is working

Bước 8. Query Log đầu tiên

Trong Grafana:

Explore

Chọn:

Loki

Query:

{namespace="todo-app"}

Bạn sẽ thấy:

2026-07-31 INFO Application started

2026-07-31 GET /api/todos

2026-07-31 ERROR Database timeout

Bước 9. Filter Backend Logs

Ví dụ:

{app="todo-backend"}

Chỉ lấy log Backend.


Bước 10. Tìm lỗi ERROR

Query:

{namespace="todo-app"}
|= "ERROR"

Ví dụ:

ERROR Connection refused

ERROR Timeout PostgreSQL

DevOps có thể nhanh chóng tìm lỗi.


Bước 11. Tìm HTTP 500

Query:

{namespace="todo-app"}
|= "500"

Ví dụ:

GET /api/users 500

Database exception

Bước 12. Debug sự cố thực tế

Giả sử Alert:

HTTP Error Rate High

Quy trình:

Bước 1

Xem Grafana Alert:

HTTP 500 tăng

Bước 2

Xem Metrics:

Request latency tăng

Bước 3

Query Logs:

ERROR

Bước 4

Tìm nguyên nhân:

Database connection timeout

Bước 13. Thử tạo lỗi

Tắt PostgreSQL

Scale:

kubectl scale deployment postgres \
--replicas=0 \
-n todo-app

Backend sẽ lỗi.


Xem Log:

ERROR

Connection refused

Khôi phục:

kubectl scale deployment postgres \
--replicas=1 \
-n todo-app

Bước 14. Kết hợp Metrics + Logs

Đây là điểm mạnh của Grafana.

Ví dụ:

Dashboard:

CPU tăng

Click:

View Logs

Xem ngay:

ERROR OutOfMemory

Một hệ thống Observability hoàn chỉnh gồm:

Metrics

+

Logs

+

Alerts

Bước 15. Debug Fluent Bit

Nếu không thấy log:

Kiểm tra Fluent Bit:

kubectl logs \
<fluent-bit-pod> \
-n logging

Kiểm tra Loki:

kubectl logs \
loki-0 \
-n logging

Kiểm tra Grafana Data Source:

Connections

↓

Loki

↓

Save & Test

Dọn dẹp

Xóa Fluent Bit:

helm uninstall fluent-bit \
-n logging

Xóa Loki:

helm uninstall loki \
-n logging

Xóa Namespace:

kubectl delete namespace logging

Những gì bạn đã học

Sau bài lab này, bạn đã biết:

  • ✅ Kubernetes Logging Architecture
  • ✅ stdout logging
  • ✅ Fluent Bit
  • ✅ Loki
  • ✅ Grafana Loki Integration
  • ✅ Query Log bằng LogQL
  • ✅ Tìm ERROR
  • ✅ Debug Production Issue
  • ✅ Kết hợp Metrics + Logs + Alerts

Bài học rút ra

Trong Production, Metrics chỉ cho chúng ta biết:

Có vấn đề xảy ra

Alerting giúp chúng ta biết:

Cần xử lý vấn đề

Nhưng Logs mới giúp trả lời:

Tại sao vấn đề xảy ra?

Một hệ thống vận hành Kubernetes hoàn chỉnh cần kết hợp:

Metrics

+

Logs

+

Alerts

Sau Lab 9, chúng ta đã hoàn thiện toàn bộ chuỗi Observability:

Prometheus

      ↓

Grafana

      ↓

Application Metrics

      ↓

Alertmanager

      ↓

Centralized Logging

Đây chính là nền tảng mà các hệ thống Kubernetes Production sử dụng để vận hành hàng trăm hoặc hàng nghìn service.


All rights reserved

Viblo
Hãy đăng ký một tài khoản Viblo để nhận được nhiều bài viết thú vị hơn.
Đăng kí