Skip to main content

Backup and restore

Only the metadata database and the etc/ directory are irreplaceable: back both up on a schedule and rehearse the restore; embedded TSDB metrics can be re-collected.

Only two things in Nightingale are genuinely irreplaceable: the metadata database and the etc/ directory. Metrics can be collected again and logs can be thrown away, but losing configuration and rules means rebuilding them by hand, one at a time. This page covers what to back up, how, and how to prove the backup actually works.

What to back up​

ContentWhereCost of not having itSuggested frequency
Metadata databaseMySQL / PostgreSQL, or n9e.dbUsers, business groups, rules, dashboards and notification config all goneDaily, keep 7–30 days
ConfigurationThe etc/ directoryEvery setting has to be worked out againOn every change (keep it in version control)
Integration templatesThe integrations/ directoryCustom dashboard and rule templates lostSame
Embedded TSDB[EmbeddedTSDB] Dir, data/tsdb by defaultHistorical metrics lostDepends how much you rely on it
Logs and evaluation records[Log] DirNo consequenceDo not back up

Evaluation records are a troubleshooting aid: they have their own retention ceiling, are cleaned up automatically, and are regenerated by every subsequent evaluation. There is nothing to preserve.

The metadata database​

Almost all the value is here.

# MySQL
mysqldump --single-transaction --routines --triggers \
-h mysql -u n9e -p n9e_v6 | gzip > n9e_v6-$(date +%F).sql.gz

# PostgreSQL
pg_dump -h postgres -U n9e -Fc n9e_v6 > n9e_v6-$(date +%F).dump

--single-transaction dumps from a consistent snapshot without locking tables.

SQLite only turns up in test environments, and it has a backup trap of its own: the data is not only in n9e.db but also in n9e.db-wal and n9e.db-shm beside it. All three must be copied and restored together — taking only n9e.db gives you a database paired with a stale WAL. Safer still, stop the process before copying.

Configuration and integration templates​

tar czf n9e-etc-$(date +%F).tar.gz etc/ integrations/

etc/ holds config.toml, the edge's etc/edge/edge.toml, metrics.yaml and the notification scripts. Better yet, keep etc/ in version control: you get a record of what changed, when and by whom, and rolling back becomes one git checkout.

config.toml contains the database password, so move credentials out before committing it — see Secret management.

The embedded TSDB​

The embedded store has no online snapshot endpoint. Keeping a copy means stopping the process and copying the directory:

systemctl stop n9e
tar czf n9e-tsdb-$(date +%F).tar.gz data/tsdb/
systemctl start n9e

Copying it while running captures half-written blocks, which may not open on restore.

Be pragmatic: if the embedded store holds "the last 15 days for looking at trends", a backup that costs downtime is not worth it — losing it is acceptable. If it is your only metric store and you cannot afford to lose it, the right answer is not to back it up but to move to an external store — see Embedded TSDB single-Center limits.

The restore procedure​

Follow the order; do not skip steps.

  1. Stop every n9e process. All of them in the cluster, n9e-edge included. A surviving instance writes heartbeats and events into the database you are restoring.

  2. Restore the database.

    gunzip < n9e_v6-2026-09-01.sql.gz | mysql -h mysql -u n9e -p n9e_v6

    Expected result: select count(*) from alert_rule; matches the count at backup time.

  3. Restore etc/ and integrations/. Confirm [DB] DSN points at the database you just restored.

  4. Restore the embedded TSDB if you backed it up and still use it: copy data/tsdb back.

  5. Start one instance. Not all of them at once. Expected result: curl --noproxy '*' http://n9e:17000/ping returns pong, and after logging in the rule list, dashboards and notification config are all there.

  6. Confirm evaluation resumed. curl --noproxy '*' http://n9e:17000/metrics | grep n9e_alert_rule_eval_total. Expected result: the count starts climbing 30 seconds after the process starts — there is a 30-second evaluation delay after boot.

  7. Start the remaining instances. Check that they all appear under System → Alerting engines.

Redis does not need restoring. It holds sessions, target heartbeat timestamps and host metadata, all of which expire; starting empty is fine. The cost is that everyone has to log in again, and target heartbeat times stay blank until the collectors next report.

Rehearsing​

A backup nobody has restored is not a backup. At minimum:

  • check the file size after every backup — a sudden drop means the dump failed part way;
  • do a full restore once a quarter into a separate test database and test instance, verifying through steps 5 and 6 that you can log in, the rules are there, and evaluation is running;
  • time the rehearsal. When it actually happens, "how long does a restore take" is a question you will have to answer.

For rebuilding from scratch in a fresh environment, see Disaster recovery.