Alert fires repeatedly or never recovers
Flapping or never-recovering alerts usually come from a threshold sitting on the data, a data gap read as recovery, a changing label set, or no recovery condition.
One rule sent forty messages overnight — fire, recover, fire, recover. Or the opposite: the problem was fixed hours ago and the event is still sitting there. These are two directions of the same machinery, so they share a page.
Decide this first: is the event flapping, or only the notification? Open the event history. A genuine alternating series of triggered/recovered rows means the event is flapping. A single event that never recovered, plus a pile of messages, means repeat notifications. The two lead to entirely different investigations.
Three timing parameters, one job each
| Parameter | Default | What it filters |
|---|---|---|
| Execution frequency | @every 60s | how often the query runs |
| For duration | 60 s | how long the condition must hold before an event exists — filters the rising edge |
| Recover duration | 0 (immediate) | how long it must be clear before recovery is declared — filters the falling edge |
The recover duration holds back the recovery event itself, not just the recovery message. Inside that window the event is still firing and still visible on the page. It is the most direct cure for flapping, and in the open-source build it is the only anti-flap lever for Prometheus rules — those have no separate recovery condition, recovery is always the inverse of the trigger, so an asymmetric "alert above 90%, recover below 70%" is not expressible.
Rule types that use the trigger-condition form (Elasticsearch, Loki, ClickHouse and the rest) do have a configurable recovery condition.
The most common cause: a data gap read as a recovery
This is the number one source of flapping and it is hard to spot.
The engine cannot tell "the value came down" from "the points for this window are missing", and it treats an empty result as a recovery in both cases. So one missed collection cycle, one dropped batch in the TSDB, or one query that times out and returns nothing produces a recovery — and the next cycle, when data returns, produces a fresh trigger.
How to confirm it is this:
- Open the rule's evaluation records, find the recovering cycle, and check that
anomaly_totalis 0 and that cycle'sseries_totalis also 0 — no series came back at all, as opposed to values dropping below the threshold; - Run a range query over the same window in the Metrics explorer and look for gaps in the line;
- Check whether
prometheus_tsdb_too_old_samples_totalon the server's/metricsis climbing (out-of-order or stale samples being discarded produce exactly these gaps).
Once confirmed, treat it from these angles:
- set the recover duration longer than one gap — with a 15-second collection interval, 60–120 seconds is a reasonable starting point;
- avoid expressions that go entirely empty when a single point is missing;
- the gap itself is an ingest problem — fix the root cause via Data source connects but queries return no data.
One counter-intuitive point worth stating plainly: a query error and an empty query result behave completely differently. On an error the whole judgement-and-recovery step for that cycle is skipped, so the event is neither refreshed nor recovered. Only a successful empty result triggers recovery. An intermittently timing-out data source therefore does not cause flapping — it causes stuck state.
The label set changed, so it became a different event
An event is identified by a hash over its label set. Change any key or value and it is a different event — the old one is recovered because it stopped appearing, the new one fires as new. What you see is endless trigger/recover on what is really the same object.
Usual sources:
- the expression carries a volatile label: container ID, pod name, an ephemeral port inside
instance; - a
relabelprocessor in a workflow attached to the rule is rewriting labels; - someone edited this host's tags in the host list while
Pushgw.LabelRewriteis on (it is by default), so the stored labels changed.
How to confirm: take an adjacent recovered/triggered pair from the history and compare their
labels key by key. Find one key that changes and you have your answer. Fix it by aggregating the
volatile label away in the expression (sum without(instance) (...)) or dropping it in a workflow.
The other direction: it never recovers
An event that will not clear, in this order:
- Is the rule still being evaluated? If the evaluation records stop at some point, no engine is running the rule any more — and without evaluation there is never a recovery. Go to Evaluation records are missing or dropped.
- Is the query erroring every cycle? Errors skip recovery entirely. Check whether
n9e_alert_rule_eval_error_total{stage="query_data"}is climbing and whethern9e_alert_eval_query_series_countis negative. - Is the recover duration too long? Until the window elapses, the event is still firing.
- Did the object disappear? Once a host is deleted from the host list, events carrying that
identare treated as muted (the evaluation record's detail readsident not exists, was muted) and will not recover on their own. - Are recovery notifications switched off? The event may have recovered on screen while you heard nothing — "notify on recovery" is a separate per-rule switch, and turning it off still records the recovery, it just does not send it.
The notifications flap, the event does not
One event in the history plus many messages is repeat notification, not flapping. Two parameters control it:
- Repeat interval (
notify_repeat_step, in minutes): how often to re-send. Setting it to 0 means one notification for the entire life of the alert. - Max notifications (
notify_max_number): 0 means unlimited.
The log states the decision for every cycle, under the prefix
alert_eval_<rule id> datasource_<ds id> event-hash-<hash>:
fired, notify_repeat_step_matched(1788440612 >= 1788438872 + 60 * 60) notify_max_number_ignore(#2 / 0)
stalled, notify_repeat_step_not_matched(1788440612 < 1788438872 + 60 * 60)
stalled, notify_repeat_step_matched(...) notify_max_number_not_matched(#5 / 5)
stalled means: the event is still firing, but this cycle sends nothing and writes no new history
row. From the user's side it simply goes quiet. The same words appear as the event stage in the
evaluation records.
When one rule carries several severities, the inhibit switch keeps only the most severe match per
object; turning it off means three messages instead of one. The inhibited counter in the evaluation
record is how many it suppressed.
Collect this before you ask
- The rule ID and its four values: execution frequency, for duration, recover duration, repeat interval;
- The event history for the flapping window, with the labels column kept — that is what makes label drift visible;
anomaly_total,series_totaland the eventstagefrom the evaluation records over the same window;- Every log line containing
event-hash-<hash>.
Redacting: replace host names, product names and container IDs inside label values, but keep the structure of which keys are changing — without that the problem is unreadable.
Next
- The full model: Evaluation, recovery and state changes
- Tuning the intervals: Evaluation interval and recovery
- Root cause of the data gaps: Data source connects but queries return no data
- Other noise-reduction options: Common noise patterns