01
Prometheus + Alertmanager
Rules live in YAML and ship through GitOps.
Rules live in the UI, scoped by business group.
Where it shines
Prometheus remains one of the best time-series foundations available, and Alertmanager’s grouping and silencing model is genuinely clean. At moderate scale, with a GitOps-comfortable team and a rule count you can hold in your head, the pair is enough and adds no components.
When Nightingale fits better
Nightingale earns its place when rules must be maintained by teams who are not SREs, when permissions need to follow business groups, when someone has to be phoned at 3am, or when the data is not all in Prometheus.
Signs it is time to switch
- Rules must be maintained by teams who are not SREs
- Someone has to be phoned at 3am, not just webhooked
- The data is no longer all in Prometheus
Migration guideFrom Prometheus + Alertmanager →
rules.yml + alertmanager.yml
- alert: MysqlReplicaLag expr: mysql_slave_lag_seconds > 30 for: 5m labels: { team: dba }route: receiver: dba-webhook group_by: [alertname]
Nightingale · alert rule
- Name
- MySQL replica lag
- Condition
- mysql_slave_lag_seconds > 30, for 5m
- Business group
- DBA
- Notify
- Phone, SMS, DingTalk
- Who can edit
- DBA group
- You
- Create in one sentence: page the DBA group by phone when mysql-prod replica lag stays above 30s for 5 minutes
- Nightingale AI
- PromQL and threshold drafted. I will not guess the business group; pick one from the list and I will create the rule.
Prometheus
VictoriaMetrics
ElasticSearch
ClickHouse
Nightingale
DingTalk
Slack