Production topology
Reference layouts for one team, one company, and multi-site with edge alerting.
Capacity planning
Load is driven by rule count, series count, event volume and database traffic; measure four numbers to know where to scale — there is deliberately no sizing table.
High availability and failure domains
Which failures a multi-instance deployment survives, which it does not, and how to test it.
Embedded TSDB single-Center limits
The embedded store lives inside one Center process: what that rules out and when to move off it.
Pushgw queue, writers and dual write
The write path: queue sizing, backpressure, multiple writers and forwarding to two backends.
Backup and restore
Only the metadata database and the etc/ directory are irreplaceable: back both up on a schedule and rehearse the restore; embedded TSDB metrics can be re-collected.
Upgrade and database migration
How schema migrations run on startup, and how to upgrade a multi-instance deployment in order.
Rollback and recovery
Undo an upgrade or a bad config push, and recover from a corrupted metadata database.
Monitor Nightingale itself
Rules that page you when the alerting system is the thing that is down.
Built-in metrics and health checks
The /metrics and health endpoints each process exposes, and the handful worth graphing.
Logs and diagnostic bundles
Four sources when Nightingale itself misbehaves: the process log, evaluation records, /metrics and pprof — collect all four before opening an issue.
Edge network partition behavior
What n9e-edge does when it loses the center: evaluation continues, notifications go local, and what syncs back later.
Disaster recovery
Rebuild from backups in a new environment, and the order in which to bring components back.