Incidents
The Incidents page (/incidents) provides the full incident management UI backed by incident-manager.
Layout
Two-panel layout:
- Left panel — incident list with a status filter and search
- Right panel — selected incident detail with timeline, runbook status, and action buttons
Incident List
Each row shows:
- Severity badge: critical (red) / high (orange) / medium (amber) / low (blue)
- Status badge: open (red) / investigating (blue) / mitigating (amber) / resolved (green)
- Title and affected service
- Age and MTTR (for resolved incidents)
- Cluster and namespace
Incident Detail
The right panel shows the full incident:
Overview section — title, description, category (node_failure / pod_crash / security / performance), affected service, and source event IDs.
Timeline — a vertical list of IncidentEvent entries. The backend records five real event types (created/acknowledged/runbook_started/escalated/resolved); the UI has a distinct icon for four of them, and falls back to a generic clock icon for escalated:
| Event type | Icon | Meaning |
|---|---|---|
| created | alert triangle | Incident opened (by system or user) |
| acknowledged | user check | Operator picked up the incident |
| runbook_started | play | Automated runbook execution began |
| resolved | check circle | Incident closed |
| escalated | (generic clock) | Severity was escalated |
Runbook execution — shows each step (kubectl / http / notify / wait) with status (pending / executing / success / failed) and output.
Post-incident review — a real, separate capability worth knowing about even though it isn't surfaced in this page's UI today: incident-manager exposes POST /api/v1/incidents/{id}/pir and GET /api/v1/incidents/{id}/pir to generate and retrieve a post-incident review for a resolved incident.
Action buttons:
[Acknowledge] → POST /api/incidents/api/v1/incidents/{id}/acknowledge
[Resolve] → POST /api/incidents/api/v1/incidents/{id}/resolve
Metrics Cards
Above the incident list, four cards sourced from GET /api/v1/incidents/metrics:
- Open Incidents — count of incidents whose status is anything other than
resolved(i.e. open + investigating + mitigating combined) - Critical — count of currently-unresolved incidents at critical severity
- Avg MTTR — mean time to resolve, averaged over incidents resolved in the last 30 days
- Total Resolved — total resolved-incident count (all-time, not scoped to the last 30 days)
Incident Correlation
incident-manager uses a 5-minute sliding window keyed on {clusterID}:{category}:{namespace} to deduplicate alert storms. If 10 pod crash events arrive in 3 minutes for the same namespace, they create one incident with 10 source events — not 10 separate incidents.
Creating Incidents Manually
POST /api/incidents/api/v1/incidents
Content-Type: application/json
{
"cluster_id": "prod-us-east",
"title": "Elevated error rate on payments-api",
"description": "5xx rate above 2% for 10+ minutes",
"severity": "high",
"category": "performance",
"affected_service": "payments-api",
"namespace": "payments"
}