Incidents & alerts
State machines with timelines, ownership, public notes, and postmortems.
Incidents and alerts are state machines, not just rows in a table: every state change closes one timeline entry and opens the next, so the history of who changed what, when, is always reconstructable.
Incidents
An incident tracks a customer-visible problem. It can be declared automatically by a monitor whose criteria say so, or manually (optionally from a reusable incident template). An incident carries:
- Severity and state — with a full state timeline.
- Affected resources — the monitors it covers.
- Owners and labels — who's responsible, how it's categorized.
- Public notes — updates that appear on status pages.
- Private notes — internal coordination.
- Postmortem — published after resolution, with its own subscriber notification.
Declaring an incident can also trigger on-call policies and subscriber notifications; resolving it flows back to the same surfaces.
Alerts
Alerts track machine-detected conditions that don't necessarily deserve a public incident. They follow the same acknowledge/resolve lifecycle, support owners, labels and notes, and can page on-call.
Related alerts group into episodes, so a flapping condition doesn't create an unbounded pile of individual alerts — episodes resolve when their underlying alerts do.
Automatic lifecycle
Monitors both open and close: when a monitor recovers, the incident or alert episode it declared is resolved automatically by the runtime — you only step in when a human decision is needed.