Skip to main content

System architecture

The core is one n9e process: no dependencies for testing, MySQL and Redis for production, several instances sharing them for a cluster, and n9e-edge for multi-site.

The architecture is small: one n9e process is the whole core. For testing it runs with no dependencies at all; production adds MySQL and Redis. Multi-site deployments with unreliable links have a separate edge mode.

What one process does​

┌──────────────────────────── n9e ────────────────────────────┐
│ Web / API rule evaluation Pushgw notification MCP │
└──┬───────────────┬──────────────┬────────────────┬──────────┘
│ │ │ │
browser / API data sources TSDB channels
Prometheus … (external or built in) DingTalk / …
│
MySQL + Redis
(metadata, coordination)

The n9e process only needs the etc/ and integrations/ directories next to the binary. Everything else is optional.

Three sizes​

Single node, testing​

Download a release, run ./n9e. Port 17000, root / root.2020.

Metadata goes into n9e.db (SQLite) next to the binary, Redis is an in-process miniredis, and metrics land in the embedded TSDB. Convenient, but not for production — see the Production readiness checklist.

Single node, production​

Edit etc/config.toml to move metadata to MySQL or PostgreSQL and use a real Redis:

[DB]
DBType = "mysql"
DSN = "root:<password>@tcp(localhost:3306)/n9e_v6?charset=utf8mb4&parseTime=True&loc=Local"

[Redis]
RedisType = "standalone"
Address = "127.0.0.1:6379"

The database is conventionally called n9e_v6 — the name has been carried since v6 and the schema scripts still use it.

Cluster​

Run n9e on several hosts with identical configuration, sharing one MySQL and one Redis. That is the whole cluster setup.

Alert rules are distributed across the instances automatically: 100 rules and 2 instances means roughly 50 each, and a rule only ever runs on one instance, so no duplicate alerts. If an instance dies, another takes over its rules.

Note that a cluster must turn off the embedded TSDB. Its data sits on one process's local disk, so with several instances each one holds a fragment.

Whether data flows through Nightingale​

These are two genuinely different shapes, and the choice decides which features you get.

Data does not flow throughData flows through
CollectionYour own Prometheus scrapesCategraf remote-writes to Nightingale
Nightingale's jobQuery the TSDBReceive, then forward per Pushgw.Writers
Host listEmptyPopulated, groupable, taggable
Self-healingUnavailableAvailable
Alert rulesFully availableFully available

Nightingale is not a long-term metric store. The embedded TSDB is the deliberate exception — see Storage.

Edge mode​

With several sites, having the central n9e query a remote site's TSDB makes alerting only as reliable as that link — and sometimes the centre cannot reach the remote store at all.

Deploy n9e-edge at the site instead: rules are still managed centrally and pushed down, but evaluation happens locally against the site's own TSDB. When the link drops, the edge keeps alerting.

See Edge data centers.