Skip to main content

High availability

Run multiple Center and Alert instances behind a load balancer, and what still needs to be a singleton.

A Nightingale cluster needs no extra component: run several n9e processes with identical config, sharing one MySQL and one Redis. This page covers doing that correctly — the load balancer in front, what the instances need from each other, and which pieces stay single even when everything else is replicated.

Building it​

Install n9e on several hosts (Binary packages) with the same configuration file. Three sections have to match:

[DB]
DSN = "n9e:<password>@tcp(mysql:3306)/n9e_v6?charset=utf8mb4&parseTime=True&loc=Local"

[Redis]
RedisType = "standalone"
Address = "redis:6379"

# only instances in the same engine cluster share rules with each other
[Alert.Heartbeat]
EngineName = "default"

Once they are up, open System → Alerting engines: every instance should be listed with a heartbeat that keeps advancing. Rules are distributed across them by consistent hash, and a rule runs on exactly one instance, so there are no duplicate alerts.

Turn the embedded TSDB off first — see "What stays a singleton" below.

Instances have to reach each other​

This step is the one that gets skipped: the instances do not only talk to the database, they call each other directly over HTTP.

Evaluation records for a rule live on the local disk of whichever instance owns that rule. When your browser lands on a different instance, that instance forwards the query to its siblings, at http://<peer ip:port>/v1/n9e/eval-records, over the service-to-service API.

So every instance in the cluster must expose that API, with the same credentials:

[HTTP.APIForService]
Enable = true
[HTTP.APIForService.BasicAuth]
n9e-cluster = "<your-own-password>"

[HTTP.APIForService] is off by default, and the user001 credentials shipped in the config file are a public example — replace them. Leaving it off does not produce an error; it just means you only ever see evaluation records for the rules the instance you reached happens to own.

An instance's identity is <IP>:<HTTP port>, and the IP is auto-detected. In containers, or on hosts with several NICs, set it explicitly so siblings do not end up with an address they cannot reach:

[Alert.Heartbeat]
IP = "n9e-1.example.internal"

Open the network path too: instance-to-instance access on port 17000.

The load balancer​

Put an L4 or L7 load balancer in front, spreading traffic across the instances' port 17000:

  • health check GET /ping, which returns pong;
  • no session affinity needed. Login state is a JWT issued and revoked through the shared Redis, so any instance accepts any request;
  • do not forward only the UI paths. The write endpoint /prometheus/v1/write, the heartbeat /v1/n9e/heartbeat, the web UI and /api/n9e/* all sit on the same port 17000 — agents and browsers come in through the same door;
  • be generous with timeouts. Dashboard and evaluation-record queries can take ten seconds or more; a 60-second backend timeout is comfortable.

The ibex RPC used by self-healing is on 20090 and is a long-lived TCP connection. Do not put it behind an L7 proxy; forward it at L4, or let agents connect directly.

What stays a singleton​

Several instances remove exactly one failure domain, the n9e process. These are not in it:

ThingSituation
Embedded TSDBData sits on one instance's local disk; a cluster must turn it off and use an external store
MySQL, RedisIf they go, everything goes — there is no degraded mode, and their HA is their own
Evaluation recordsOn each instance's local disk, with no replica; losing the instance loses its records
DingTalk stream modeRuns on the leader only — the lowest-sorted IP:port among live instances, elected automatically

Turning the embedded store off:

[EmbeddedTSDB]
Enable = false

For the migration path see External TSDB and dual-write migration.

About n9e-alert and n9e-pushgw​

Rule evaluation and ingestion already live inside n9e, so scaling out means running more n9e processes. Splitting them into separate processes is optional, and neither binary is in the release package — the Makefile has build-alert and build-pushgw targets, so a split deployment has to be built from source.

Next​