Skip to main content

Built-in metrics

What each Nightingale process exposes about itself on /metrics.

Every process exposes /metrics on its own HTTP port, in Prometheus format:

curl http://n9e:17000/metrics

Controlled by [HTTP] ExposeMetrics, which is on by default.

An alerting system that is down will not alert you about being down, so these should be scraped by your monitoring and have rules on them — see Monitor Nightingale itself.

The handful to watch first​

Ordered by how early they show a problem:

MetricWhy
n9e_alert_rule_eval_total / n9e_alert_rule_eval_error_totalHow many evaluations ran and how many failed. Evaluation stopping shows here first
n9e_alert_rule_eval_duration_msTime per evaluation. A slowing data source surfaces here
n9e_alert_alert_queue_sizeEvent queue backlog. Sustained growth means the downstream cannot keep up
n9e_alert_query_data_error_totalFailed data source queries
n9e_center_http_request_duration_secondsWeb and API latency
n9e_db_operation_latency_secondsMetadata database latency

Full list​

Nightingale exposes 55 metrics of its own on every process's /metrics endpoint ([HTTP] ExposeMetrics is on by default).

Go runtime and process metrics are left out below — every Go program has those, and they say nothing about Nightingale.

The 16 marked † come from the source rather than from a live scrape: a labelled metric does not appear on /metrics until something first observes it, so a freshly started instance will not show them. That is expected — write rules against the name regardless.

n9e_alert_*​

MetricTypeMeaning
n9e_alert_alert_notify_error_total †counterNumber of send msg.
n9e_alert_alert_notify_total †counterNumber of send msg.
n9e_alert_alert_queue_sizegaugeThe size of alert queue.
n9e_alert_alerts_totalcounterTotal number alert events.
n9e_alert_eval_log_drop_totalcounterNumber of eval log records dropped (full write queue, or still oversized after full degradation).
n9e_alert_eval_log_query_reject_totalcounterNumber of eval log queries rejected by the concurrency gate.
n9e_alert_eval_query_series_countgaugeNumber of series retrieved from data source after query.
n9e_alert_heartbeat_error_count †counterNumber of heartbeat error.
n9e_alert_mute_total †counterNumber of mute.
n9e_alert_notify_record_queue_sizegaugeThe size of notify record queue.
n9e_alert_query_data_error_total †counterNumber of rule eval query data error.
n9e_alert_query_data_totalcounterNumber of rule eval query data.
n9e_alert_record_eval_duration_ms †gaugeDuration of record eval in milliseconds.
n9e_alert_record_eval_error_total †counterNumber of record eval errors, labeled by stage.
n9e_alert_record_eval_series_count †gaugeSeries count produced by the latest record eval; negative values encode error states.
n9e_alert_record_eval_total †counterNumber of record eval.
n9e_alert_rule_eval_duration_msgaugeDuration of rule eval in milliseconds.
n9e_alert_rule_eval_error_total †counterNumber of rule eval error.
n9e_alert_rule_eval_totalcounterNumber of rule eval.
n9e_alert_sub_event_total †counterNumber of sub event.
n9e_alert_var_filling_query_total †counterNumber of var filling query.

n9e_center_*​

MetricTypeMeaning
n9e_center_http_request_duration_secondshistogramHTTP request latencies in seconds.
n9e_center_redis_operation_latency_secondshistogramHistogram of latencies for Redis operations
n9e_center_uptimecounterHTTP service uptime.

n9e_cron_*​

MetricTypeMeaning
n9e_cron_durationgaugeCron method use duration, unit: ms.
n9e_cron_sync_numbergaugeCron sync number.

n9e_db_*​

MetricTypeMeaning
n9e_db_operation_latency_secondshistogramHistogram of latencies for DB operations
n9e_db_operation_totalcounterTotal number of DB operations
n9e_db_pool_idle_connectionsgaugeThe number of idle connections
n9e_db_pool_in_use_connectionsgaugeThe number of connections currently in use
n9e_db_pool_max_idle_closed_totalcounterThe total number of connections closed due to SetMaxIdleConns
n9e_db_pool_max_idle_time_closed_totalcounterThe total number of connections closed due to SetConnMaxIdleTime
n9e_db_pool_max_lifetime_closed_totalcounterThe total number of connections closed due to SetConnMaxLifetime
n9e_db_pool_max_open_connectionsgaugeMaximum number of open connections to the database
n9e_db_pool_open_connectionsgaugeThe number of established connections both in use and idle
n9e_db_pool_wait_count_totalcounterThe total number of connections waited for
n9e_db_pool_wait_duration_seconds_totalcounterThe total time blocked waiting for a new connection

n9e_pushgw_*​

MetricTypeMeaning
n9e_pushgw_drop_sample_totalcounterNumber of drop sample.
n9e_pushgw_forward_duration_secondshistogramForward samples to TSDB. latencies in seconds.
n9e_pushgw_http_request_duration_secondshistogramHTTP request latencies in seconds.
n9e_pushgw_proxy_forward_error_total †counterNumber of forward errors on /proxy/v1/write.
n9e_pushgw_proxy_forward_total †counterNumber of forwards performed by /proxy/v1/write.
n9e_pushgw_proxy_remote_write_body_too_large_totalcounterNumber of /proxy/v1/write requests rejected with 413 due to body size over limit.
n9e_pushgw_proxy_remote_write_inflightgaugeCurrent number of in-flight requests on /proxy/v1/write.
n9e_pushgw_proxy_remote_write_over_limit_totalcounterNumber of /proxy/v1/write requests rejected with 429 due to in-flight over limit.
n9e_pushgw_proxy_remote_write_totalcounterNumber of /proxy/v1/write requests received.
n9e_pushgw_push_queue_error_total †counterNumber of push queue error.
n9e_pushgw_push_queue_over_limit_error_totalcounterNumber of push queue over limit.
n9e_pushgw_redis_operation_latency_secondshistogramHistogram of latencies for Redis operations
n9e_pushgw_sample_queue_sizegaugeThe size of sample queue.
n9e_pushgw_sample_received_by_identcounterNumber of sample push by ident.
n9e_pushgw_samples_received_totalcounterTotal number samples received.
n9e_pushgw_write_error_total †counterNumber of write error.
n9e_pushgw_write_totalcounterNumber of write.

n9e_sandbox_*​

MetricTypeMeaning
n9e_sandbox_exec_activegaugeCurrently running sandbox executions.