Skip to main content

Built-in metrics and health checks

The /metrics and health endpoints each process exposes, and the handful worth graphing.

n9e, n9e-edge, n9e-alert and n9e-pushgw share one HTTP skeleton, so the endpoints below are the same on every process — only the port differs (17000 for Center, 19000 for edge).

Endpoints every process has​

EndpointReturnsUse it for
GET /pingpongLoad balancer and Kubernetes probes; the cheapest check
GET /pidThe process PIDConfirming which process you reached
GET /ppidThe parent PIDTelling whether a supervisor started it
GET /addrThe caller's addressAnswering "which IP am I actually coming from"
GET /api/n9e/versionThe version, e.g. v9.1.1Checking instances one by one during a rolling upgrade
GET /metricsPrometheus-format self metricsSee the next section
GET /api/debug/pprof/*Go pprofCPU and memory work — see Logs and diagnostic bundles
GET /dumper/syncConfig sync status per cacheLocal requests only; answers "I changed config and nothing happened"

Try it:

curl --noproxy '*' http://n9e:17000/ping # pong
curl --noproxy '*' http://n9e:17000/api/n9e/version # v9.1.1

/ping only proves the HTTP server is still listening. It says nothing about whether the database is reachable or rules are being evaluated — that judgement comes from /metrics.

Two switches: [HTTP] ExposeMetrics defaults to true, and [HTTP] PProf defaults to true in etc/config.toml but false in etc/edge/edge.toml.

The handful worth graphing​

/metrics is thousands of lines, most of it Go runtime internals and per-route HTTP histograms. These are the ones that answer "is the system still doing its job".

Is evaluation still running

MetricMeaning
n9e_center_uptimeSeconds since process start; a reset to zero means a restart (Center only)
n9e_alert_rule_eval_totalTotal rule evaluations, no labels. If it stops climbing, evaluation stopped
n9e_alert_rule_eval_error_totalEvaluation errors, labelled datasource / stage / busi_group / rule_id
n9e_alert_rule_eval_duration_msLast evaluation duration per rule, labelled rule_id / datasource_id
n9e_alert_query_data_error_totalFailed data source queries, labelled datasource
n9e_alert_alerts_totalAlert events produced, labelled cluster / type / busi_group
n9e_alert_alert_queue_sizeIn-memory event backlog; persistently non-zero means processing is behind

Is data still arriving

MetricMeaning
n9e_pushgw_samples_received_totalSamples received, labelled channel (prometheus, opentsdb, …)
n9e_pushgw_sample_received_by_identSamples per host, labelled host_ident; makes a silent host obvious
n9e_pushgw_sample_queue_sizeForwarding queue depth, labelled queueid (one queue per CPU)
n9e_pushgw_push_queue_over_limit_error_totalWhole remote-write requests rejected because the queues were over the watermark
n9e_pushgw_push_queue_error_totalSamples dropped because one queue was full, labelled queueid
n9e_pushgw_drop_sample_totalSamples discarded by the [Pushgw.DebugSample]-style drop filters
n9e_pushgw_write_total / n9e_pushgw_forward_duration_secondsSamples written out and how long it took, labelled url

Are the dependencies healthy

MetricMeaning
n9e_db_pool_open_connections / n9e_db_pool_in_use_connectionsPool usage; approaching [DB] MaxOpenConns means queuing
n9e_db_pool_wait_count_total / n9e_db_pool_wait_duration_seconds_totalWaits for a connection; non-zero means tune the pool or find the slow queries
n9e_db_operation_totalDatabase operations, labelled operation / status / table
n9e_center_redis_operation_latency_secondsRedis latency histogram
n9e_center_http_request_duration_secondsHTTP latency histogram, labelled code / method / path

The embedded TSDB (only on a Center with [EmbeddedTSDB] on)

It uses the Prometheus storage engine, so its metrics are the standard ones: prometheus_tsdb_head_series (active series right now), prometheus_tsdb_head_samples_appended_total, prometheus_tsdb_storage_blocks_bytes, prometheus_tsdb_wal_corruptions_total, and prometheus_engine_queries (queries executing or queued).

The alerting engine heartbeat​

There is a more direct view than metrics: System → Alerting engines. Every instance writes a heartbeat once a second, and the page lists the live instances grouped by engine cluster.

A Last heartbeat frozen at some point means that instance is in trouble; a row that disappears entirely means it went over 30 seconds without a heartbeat and was removed from the hash ring.

Version information lives under System → About: frontend version, backend version, and which features each data source type supports.

The about pageThe about page

Labelled counters do not exist until they fire​

A Prometheus CounterVec only appears on /metrics once some label combination has been recorded. On a freshly started instance, n9e_alert_rule_eval_error_total and n9e_alert_query_data_error_total are therefore absent, not zero.

Two ways that bites:

  • using absent(n9e_alert_rule_eval_error_total) to mean "are there errors" — it fires permanently while everything is fine;
  • thresholding an error counter without rate() — it steps from nothing to something the first time it appears.

The reliable approach is to judge on a metric that is always present, such as the growth rate of n9e_alert_rule_eval_total. Concrete rules are in Monitor Nightingale itself.

Closing these endpoints down​

Neither /metrics nor /api/debug/pprof/* requires authentication, and they share a port with the web UI and the ingest endpoints. In production:

  • set [HTTP] PProf = false and turn it on temporarily when you need a profile;
  • let only your internal collector reach /metrics, allowlisted by path at the gateway;
  • /dumper/sync already refuses non-local requests, so it needs nothing extra.