从零搭一套服务器监控告警:Prometheus + Grafana + 飞书通知
从零开始,把 Prometheus、Grafana、Alertmanager 和节点采集器搭起来,最后把告警收到飞书群里。全程 Docker,一台 4 核 8G 的机器就够。
先看清数据是怎么流的
node-exporter (9100) ─┐
cadvisor (8080) ──────┼─→ Prometheus (9090) ──→ Grafana (3000) 看板
你的应用 /metrics ────┘ │
└──→ Alertmanager (9093) ──→ 飞书 / 邮件
几个容易搞混的点:
- Prometheus 是拉模式(pull),它主动去所有目标抓指标。所以你得提前把目标写进配置,或者用服务发现。
- 每个要监控的机器上跑一个 node-exporter,它把系统指标暴露在
/metrics。 - Alertmanager 单独跑,Prometheus 算出告警后推给它,由它负责分组、去重、路由。
目录结构
先把文件位置定下来,配置全部挂载进去,改配置不用重建镜像:
monitoring/
├── compose.yaml
├── prometheus/
│ ├── prometheus.yml
│ └── rules/
│ └── node.yml
├── alertmanager/
│ └── alertmanager.yml
└── grafana/
└── provisioning/
└── datasources/
└── prometheus.yml
第一步:写 compose.yaml
版本就选当前 LTS 线。Prometheus 3.x 已经原生支持 OTLP 接入,和 OpenTelemetry 的协议对接问题解决了;Grafana 13 支持把告警直接推到飞书和钉钉,不用自己写转发服务。
services:
prometheus:
image: prom/prometheus:v3.13.1
container_name: prometheus
ports:
- "9090:9090"
volumes:
- ./prometheus/prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./prometheus/rules:/etc/prometheus/rules:ro
- prom_data:/prometheus
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.path=/prometheus
- --storage.tsdb.retention.time=30d
- --web.enable-lifecycle # 支持热重载配置
restart: unless-stopped
node-exporter:
image: prom/node-exporter:latest
container_name: node-exporter
pid: host # 不加这个,看不到宿主机全部进程
volumes:
- /proc:/host/proc:ro
- /sys:/host/sys:ro
- /:/rootfs:ro
command:
- --path.procfs=/host/proc
- --path.sysfs=/host/sys
- --path.rootfs=/rootfs
- --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)
restart: unless-stopped
alertmanager:
image: prom/alertmanager:v0.29.0
container_name: alertmanager
ports:
- "9093:9093"
volumes:
- ./alertmanager/alertmanager.yml:/etc/alertmanager/alertmanager.yml:ro
- am_data:/alertmanager
restart: unless-stopped
grafana:
image: grafana/grafana:13.0.2
container_name: grafana
ports:
- "3000:3000"
environment:
GF_SECURITY_ADMIN_PASSWORD: change-me-now
GF_USERS_ALLOW_SIGN_UP: "false"
volumes:
- ./grafana/provisioning:/etc/grafana/provisioning:ro
- grafana_data:/var/lib/grafana
depends_on:
- prometheus
restart: unless-stopped
volumes:
prom_data:
am_data:
grafana_data:
node-exporter 那段有三个参数不加就取不到数据,我列一下:pid: host、--path.procfs、--path.rootfs。少了任何一个,你会看到指标是空的或者只反映容器自己。
第二步:Prometheus 配置
# prometheus/prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
external_labels:
cluster: prod-1
alerting:
alertmanagers:
- static_configs:
- targets: ["alertmanager:9093"]
rule_files:
- /etc/prometheus/rules/*.yml
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ["localhost:9090"]
- job_name: node
static_configs:
- targets: ["node-exporter:9100"]
# 监控更多机器就这样加
# - targets: ["10.0.1.11:9100", "10.0.1.12:9100"]
- job_name: app
metrics_path: /metrics
scrape_interval: 30s
static_configs:
- targets: ["app:8000"]
scrape_interval 怎么定?默认 15 秒。核心业务想看得更细可以降到 5 秒,但存储压力是线性涨的——采样频率翻三倍,磁盘占用也差不多翻三倍。非核心的 30 秒或 60 秒就够。
第三步:写告警规则
# prometheus/rules/node.yml
groups:
- name: host
rules:
- alert: InstanceDown
expr: up == 0
for: 2m
labels:
severity: critical
annotations:
summary: "实例 {{ $labels.instance }} 已下线"
description: "已持续 2 分钟无法抓取指标。"
- alert: HighCPU
expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 85
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} CPU 持续高于 85%"
- alert: LowDiskSpace
expr: |
(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}
/ node_filesystem_size_bytes{fstype!~"tmpfs|overlay"}) * 100 < 15
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} 磁盘剩余不足 15%"
- alert: HighMemory
expr: |
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 > 90
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} 内存使用率超过 90%"
两个细节值得说。内存告警一定用 MemAvailable,不要用 MemFree。 前者已经把可回收的 page cache 算进去了,后者会把正常运行的系统报成 99% 占用——我见过太多人被这个指标坑。磁盘告警用剩余百分比而不是绝对值,因为一台 2T 的盘剩 50G 很危险,一台 100G 的盘剩 50G 还很宽裕。
for: 10m 是防抖用的。指标瞬间抖动不算,要持续 10 分钟才发告警。没有这个,半夜告警能把你手机震醒八次。
第四步:接到飞书
Alertmanager 本身不支持飞书,但飞书群机器人接受一个简单的 JSON webhook,配一下就通。
先在飞书群里加个"自定义机器人",拿到 webhook 地址,然后:
# alertmanager/alertmanager.yml
global:
resolve_timeout: 5m
route:
group_by: ["alertname", "instance"]
group_wait: 10s # 等 10 秒看有没有同类告警,一起发
group_interval: 5m
repeat_interval: 4h # 同一个告警没恢复,4 小时提醒一次
receiver: feishu
receivers:
- name: feishu
webhook_configs:
- url: "http://feishu-bridge:8080/alert"
send_resolved: true
飞书的 payload 格式和 Alertmanager 默认的不一样,需要一层转换。最简单的方法是加一个轻量 bridge,把 Alertmanager 的 JSON 转成飞书的 {"msg_type":"text","content":{"text":"..."}}。写个二十行的 Flask 服务就够,用模板渲染出标题和描述。
如果你不想自己写,也可以直接升级到 Grafana 13 —— 它在 2026 年已经内置了飞书和钉钉的告警通道,把 Prometheus 作为数据源接进去,告警直接在 Grafana 里配就行,省掉这一层。
第五步:Grafana 初始化
别用界面点,用 provisioning 文件,配置就能进 Git:
# grafana/provisioning/datasources/prometheus.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
editable: false
启动后浏览器打开 http://localhost:3000,默认账号密码都是 admin(密码我们已经在环境变量里改掉了)。然后直接导入社区模板 Node Exporter Full(ID:1860),主机监控的 CPU、内存、磁盘、网络就全有了,不用自己拖面板。
第六步:验证一遍
docker compose up -d
docker compose ps
# 看 Prometheus 有没有抓到目标
curl -s localhost:9090/api/v1/targets | head -c 500
# 手动触发一次告警,确认通知链路通
curl -H "Content-Type: application/json" -d '[{"labels":{"alertname":"TestAlert","severity":"warning","instance":"test"}}]' \
http://localhost:9093/api/v2/alerts
第二条命令发出去,飞书群里应该几秒内就有消息。这一步一定要做,很多人配完等到真出事才发现通知根本没通。
改完 Prometheus 配置不想重启容器的话:
curl -X POST http://localhost:9090/-/reload # 前提是开了 --web.enable-lifecycle
第七步:把告警收敛住
真出事的时候,最怕的不是没告警,是告警刷屏。一台机器磁盘满了,可能同时触发磁盘、写入延迟、服务异常、连接数超限七八条告警,电话被打爆,但根因只有一个。
Alertmanager 有三个机制专门治这个,都要配上。
分组(group_by):把同一类告警合并成一条通知。我们已经配了 group_by: ["alertname", "instance"],再配合 group_wait: 10s,10 秒内的同类告警会合成一条发出去。
抑制(inhibit_rules):高优先级告警出现时,压掉它引起的次生告警。典型场景是"机器下线"和"机器上的服务不可用"——后者是前者的结果,不用重复告警:
inhibit_rules:
- source_match:
alertname: InstanceDown
target_match_re:
alertname: "HighCPU|HighMemory|HighDiskUsage"
equal: ["instance"]
意思是:同一台机器上,只要 InstanceDown 在响,就不发它的 CPU/内存/磁盘告警。
静默(silence):计划内维护时手动关掉告警。这个在 Alertmanager 的 Web 界面(9093 端口)上点几下就行,不需要改配置:
http://localhost:9093/#/silences
维护窗口开始前设一个 silence,结束自动失效。比临时注释掉规则文件安全得多。
告警分级也很重要。我在规则里用了 severity: critical 和 severity: warning 两级,路由可以按级别分:
critical:直接进电话/短信(PagerDuty、飞书加急)warning:进群消息,工作时间看就行
route:
receiver: feishu
routes:
- matchers: [severity="critical"]
receiver: pager
repeat_interval: 30m
- matchers: [severity="warning"]
receiver: feishu
repeat_interval: 4h
规则很简单但很有效:只有真的需要人立刻起床的,才进 critical。 判断标准是"这条告警值不值得凌晨三点把人叫起来",不值得的一律 warning。
几个容易翻车的点
1. 容器里的 localhost。 prometheus.yml 里写 localhost:9100 是指 Prometheus 容器自己,抓不到 node-exporter。服务之间用服务名互访,也就是 node-exporter:9100。这个错误新手基本都会犯一次。
2. 磁盘被撑爆。 默认保留 15 天,如果采集目标多、指标基数高,磁盘涨得很快。要么调 --storage.tsdb.retention.time,要么上 VictoriaMetrics——它的压缩比能省 70% 以上的磁盘,而且兼容 PromQL。
3. 指标基数爆炸。 别把用户 ID、请求 ID 这种高基数的东西当标签(label)。一张报表带几万个不同 label 组合,Prometheus 内存会直接飙起来。Prometheus 3.x 可以配 sample_limit 和 label_limit 兜底。
4. 时区。 容器默认 UTC,Grafana 面板时间会差 8 小时。要么给 Grafana 设 GF_DATE_FORMATS_DEFAULT_TIMEZONE=Asia/Shanghai,要么在面板里手动改。
最后
自建监控真正的成本不在搭建,在告警质量。搭起来两小时,把告警调到一个"该响的响、不该响的不响"的状态,得花几周。
所以我的建议是:先只配四条告警(实例下线、CPU、内存、磁盘),跑两周,看哪些是误报、哪些是真事,再慢慢加。 一上来配三十条规则,结果就是所有人开始无视告警——那比不监控还危险。
DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。
更多推荐



所有评论(0)