Skip to main content
Trang này dùng cho vận hành VMS master, VMS agent collector, dashboard, alert và synthetic/readiness check. Mục tiêu là triage nhanh theo tầng Infrastructure -> Service -> User -> Business Flow, sau đó route đúng owner.

Quy tắc triage nhanh

1

Xác định scope ảnh hưởng

Kiểm tra issue chỉ ảnh hưởng một host/service, một collector, một dashboard hay toàn bộ VMS master.
2

Kiểm tra freshness

So sánh thời điểm metric cuối, heartbeat cuối, thời điểm alert phát sinh và maintenance window hiện tại.
3

Khoanh vùng layer

Phân loại issue thuộc Infrastructure, Service, User check hay Business Flow để tránh route sai owner.
4

Đối chiếu inventory

Kiểm tra tag system, environment, service, owner, criticality, scope trước khi kết luận mất dữ liệu.
5

Escalate có bằng chứng

Khi cần escalate, gửi kèm collector id, host, service, dashboard, alert id, timestamp và log liên quan.

Collector không gửi dữ liệu

Dashboard stale hoặc thiếu dữ liệu

Alert noise

Service health check fail

Pre-market readiness fail

Khi cần gửi thông tin hỗ trợ