Skip to main content

Retries and delivery status

Failed sends are retried only for connection failures, with parameters per send type; the outcome of every message is in the notification log, kept for seven days.

Where this page ends: you know where a notification can get stuck between "decided to send" and "they received it", which failures are retried and which are not, and where to look afterwards to find out what happened to one specific message.

How a message actually leaves​

The send path branches on the media type's request type (request_type):

Request typeHow it is sentConcurrency control
http (DingTalk, WeCom, Feishu Card, Telegram, Callback, your own webhooks)Queued per media type, sent asynchronouslyConcurrency on the media type
smtp (Email)Sent synchronously through the mail connection poolBatch on the media type (mails per connection)
flashduty / pagerduty / scriptCalled directly, synchronouslynone

When the HTTP queue is full the message is dropped with a failure record reading failed to enqueue notify task, queue is full. Seeing that means the target is too slow or the concurrency is set too low.

Retry only covers "could not connect"​

This is the part people misread. The retry loop for HTTP media types works like this:

  • The request never went out (DNS failure, connection refused, timeout) → wait one retry interval and try again, up to retry times attempts. When all of them fail the record reads all retries failed, last error: <the last error>.
  • The target returned any HTTP response → return immediately, no retry. 200 counts as success, anything else as failure, and the record reads status_code:<code>, response:<body>.

So DingTalk answering 400 "keywords not in content", or WeCom answering 45009 "rate limited", is not retried. That is deliberate: re-sending something the target explicitly rejected achieves nothing and is unfriendly to downstream deduplication.

One edge worth knowing: retry times set to 0 means nothing is sent at all — the loop body never runs, and you get an all retries failed record with an empty error. Do not use it as a throttle.

Retry settings differ per request type​

Settinghttpflashduty / pagerduty
TimeoutTimeout (ms), 10000 on the built-insTimeout (ms), 5000 on the built-in FlashDuty
RetriesRetry times, 3 on the built-insRetry times, 3 when unset
WaitRetry interval (ms), 100 on the built-insRetry sleep (ms), 1000 when unset
Attempts actually madeequals retry timesretry times + 1

The defaults the "add media type" form pre-fills are not identical to the built-ins (timeout 10000, concurrency 3, retries 3, interval 3000). Adjust them to taste.

The three paths also disagree about what an error status code means:

  • http — only 200 is success; anything else is recorded as a failure but is not retried;
  • pagerduty — 200 and 202 are success; any other status is retried up to the retry count;
  • flashduty — as long as the request left the machine it counts as success; the status code is recorded but never judged, because re-sending would break the target's deduplication.

Dropped before anything is sent​

In these cases no request is made at all, but a failure record is still written:

What the record saysCauseFix
message_template not foundThat notification config has no message templatePick one in the notification rule. FlashDuty, PagerDuty and Callback media types need none, which is why the UI hides the field for them
notify_channel not foundThe media type was deleted — or disabled; the engine cache holds only enabled onesRe-enable it, or switch to another
failed to enqueue notify task, queue is fullThe HTTP send queue backed upRaise the media type's concurrency, or find out why the target is slow

Two other cases produce no record at all, because the event never reached the notification stage: an event pipeline on the notification rule dropped it, or it is a recovery event and the rule does not notify on recovery.

Where to see whether it landed​

The open-source edition has no standalone "notification records" menu page; records hang off the event: Alerts & Notifications → Events → open an event → Notification records → View detail.

The drawer shows Alert rule notification and Subscription rule notification as two tables, with columns notification rule ID, channel, username, target and status. Three things to keep in mind:

  • Several records for the same (channel, target) pair are merged into one row. The worse status wins and the detail strings are concatenated, so a single row can carry both a success and a failure fragment.
  • Targets are masked by default, last 8 characters replaced with asterisks. Only email, SMS, voice, script and mute records show the full value — for those the target is something the recipient can see anyway.
  • Records are written asynchronously (an in-memory queue flushed in batches every 100 ms), so refreshing the instant after a send may show nothing yet.

There are three statuses: success, failure, and muted. The last one comes from a "mute notification only" mute rule: the channel shows as mute, the target as id=<mute rule id>, and the detail names the mute rule. That is how you answer "what exactly did we swallow during that window" — see Mute rules.

Records are kept for 7 days​

notification_record is a large, continuously written table. The center process cleans it once a day at 01:00 and keeps 7 days by default. To change that, add this to the [Center] section of the center's config file:

[Center]
# retention days for notification records, default 7
CleanNotifyRecordDay = 30

If you need long-term retention, export the rows elsewhere rather than raising this value — the table grows in proportion to your alert volume.

Next​