Data source connects but queries return no data
A source passes its test but returns nothing: check write, store and query in turn — nothing written, time range past the last point, label mismatch, wrong store.
Categraf target is offline
A target shows offline when its heartbeat has not reached Redis in time — clock skew, a wrong heartbeat URL or network policy; update time is not a liveness signal.
Query works but the alert does not fire
Query has data but nothing fires: the evaluation records hold the answer — for-duration not met, wrong datasource_queries, group or window mismatch, disabled, muted.
Evaluation records are missing or dropped
Missing evaluation records mean one of four things: not recorded, not evaluated, query rejected or cycle skipped; overload and datasource timeouts are the usual causes.
Alert fires repeatedly or never recovers
Flapping or never-recovering alerts usually come from a threshold sitting on the data, a data gap read as recovery, a changing label set, or no recovery condition.
Alert event exists but no notification arrives
Event but no message: in the notification log, no record means a mute or notification rule stopped it; a failure points at the channel, a success at the receiver.
Notification template rendering fails
A message body that is error text, or a field that is empty, is a template rendering problem that still logs as success; preview against a real event to catch it.
Event pipeline routes to the wrong destination
Wrong recipients usually mean the destination was decided earlier than expected: processor order, a label rewrite before the match, or the two mount points' scopes.
Workflow Try run succeeds but production execution fails
Try run passes, production fails: real events may lack labels the mock had, and the alert process — not the UI — runs the workflow with its own permissions and network.
Embedded TSDB has fragmented or missing data
Gaps in the embedded TSDB rarely raise errors — retention, disk pressure or restarts; three timestamp metrics show whether the gap is on the write or the query side.
Edge loses contact with the center
During a partition the edge keeps alerting locally but its events never reach the center — that is expected; an edge that never connected is a misconfiguration.
MCP authentication fails
MCP authentication fails at the gateway, the token check or OAuth — the status code says which; usual causes: wrong header name, expired token, bad redirect URL.
MCP tools are missing or denied
"No such tool" has three shapes — write tools not enabled, toolset disabled, or the token's user lacks the permission — and the tools/list response tells them apart.
Upgrade or database migration fails
Unknown column errors, blank pages or a hung process after an upgrade mean the schema did not fully migrate — usually missing database privileges or an interrupted run.
Performance and backlog symptoms
Late alerts, growing queues and high database CPU each map to one built-in metric that localises the cause; pick the metric by symptom, then work back to the root.