Skip to main content
Version: 1.0

Incidents

The Incidents page (/incidents) provides the full incident management UI backed by incident-manager.

Layout​

Two-panel layout:

  • Left panel — incident list with a status filter and search
  • Right panel — selected incident detail with timeline, runbook status, and action buttons

Incident List​

Each row shows:

  • Severity badge: critical (red) / high (orange) / medium (amber) / low (blue)
  • Status badge: open (red) / investigating (blue) / mitigating (amber) / resolved (green)
  • Title and affected service
  • Age and MTTR (for resolved incidents)
  • Cluster and namespace

Incident Detail​

The right panel shows the full incident:

Overview section — title, description, category (node_failure / pod_crash / security / performance), affected service, and source event IDs.

Timeline — a vertical list of IncidentEvent entries. The backend records five real event types (created/acknowledged/runbook_started/escalated/resolved); the UI has a distinct icon for four of them, and falls back to a generic clock icon for escalated:

Event typeIconMeaning
createdalert triangleIncident opened (by system or user)
acknowledgeduser checkOperator picked up the incident
runbook_startedplayAutomated runbook execution began
resolvedcheck circleIncident closed
escalated(generic clock)Severity was escalated

Runbook execution — shows each step (kubectl / http / notify / wait) with status (pending / executing / success / failed) and output.

Post-incident review — a real, separate capability worth knowing about even though it isn't surfaced in this page's UI today: incident-manager exposes POST /api/v1/incidents/{id}/pir and GET /api/v1/incidents/{id}/pir to generate and retrieve a post-incident review for a resolved incident.

Action buttons:

[Acknowledge]  →  POST /api/incidents/api/v1/incidents/{id}/acknowledge
[Resolve] → POST /api/incidents/api/v1/incidents/{id}/resolve

Metrics Cards​

Above the incident list, four cards sourced from GET /api/v1/incidents/metrics:

  • Open Incidents — count of incidents whose status is anything other than resolved (i.e. open + investigating + mitigating combined)
  • Critical — count of currently-unresolved incidents at critical severity
  • Avg MTTR — mean time to resolve, averaged over incidents resolved in the last 30 days
  • Total Resolved — total resolved-incident count (all-time, not scoped to the last 30 days)

Incident Correlation​

incident-manager uses a 5-minute sliding window keyed on {clusterID}:{category}:{namespace} to deduplicate alert storms. If 10 pod crash events arrive in 3 minutes for the same namespace, they create one incident with 10 source events — not 10 separate incidents.

Creating Incidents Manually​

POST /api/incidents/api/v1/incidents
Content-Type: application/json

{
"cluster_id": "prod-us-east",
"title": "Elevated error rate on payments-api",
"description": "5xx rate above 2% for 10+ minutes",
"severity": "high",
"category": "performance",
"affected_service": "payments-api",
"namespace": "payments"
}