Skip to main content
Version: 2.0

Incidents

The Incidents page (/incidents) is your incident response workspace, backed by the incident manager. KubeOpera opens incidents automatically from anomalies, self-healing events and security findings — and you can open them by hand.

Layout​

  • Left: the incident list, with status and severity filters and search.
  • Right: the selected incident — overview, timeline, runbooks, actions and review.

Four cards across the top summarize your reliability:

CardMeaning
Open IncidentsIncidents not yet resolved (open, investigating or mitigating).
CriticalUnresolved incidents at critical severity.
Avg MTTRMean time to resolve, over incidents resolved in the last 30 days.
Total ResolvedAll incidents resolved to date.

The incident list​

Each row shows:

  • Severity: critical (red), high (orange), medium (amber), low (blue).
  • Status: open, investigating, mitigating, resolved.
  • The title and affected service, cluster and namespace.
  • Age, and time to resolve once resolved.

One incident, not a hundred alerts​

When something breaks, it rarely breaks once. The incident manager groups events that share the same cluster, category and namespace within a five-minute window into a single incident. Ten pod crashes in three minutes in the payments namespace become one incident with ten source events — not ten pages to your on-call engineer.

Working an incident​

Overview​

Title, description, category (node_failure, pod_crash, security, performance), affected service, and the source events that opened it.

Timeline​

Everything that happens to the incident is recorded:

EventMeaning
createdThe incident was opened, by KubeOpera or by a person.
acknowledgedSomeone took ownership.
runbook_startedAn automated runbook began.
escalatedSeverity was raised or the incident was escalated to another responder.
resolvedThe incident was closed.

Runbooks​

Runbooks are automated response steps that run when an incident matches. Each step is one of kubectl, http, notify or wait, and shows its status (pending, executing, success, failed) and output as it runs. See Incident manager to write your own.

Actions​

ButtonWhat it does
AcknowledgeTake ownership; notifications for this incident stop escalating.
EscalateRaise severity or hand the incident to the next responder.
Investigate with AILaunch an Incident Responder agent with this incident as context.
ResolveClose the incident and record the time to resolve.

Post-incident review​

When an incident is resolved, select Generate review to create a post-incident review (PIR): a summary of what happened, the timeline, contributing factors, what was done, and follow-up actions. Edit it, then share it with your team. Reviews stay attached to the incident for future reference.

Open an incident yourself​

Select New Incident, or use the API:

POST /api/incidents/api/v1/incidents
Content-Type: application/json

{
"cluster_id": "prod-us-east",
"title": "Elevated error rate on payments-api",
"description": "5xx rate above 2% for 10+ minutes",
"severity": "high",
"category": "performance",
"affected_service": "payments-api",
"namespace": "payments"
}

Next steps​