Incidents
The Incidents page (/incidents) is your incident response workspace, backed by the incident manager. KubeOpera opens incidents automatically from anomalies, self-healing events and security findings — and you can open them by hand.
Layout
- Left: the incident list, with status and severity filters and search.
- Right: the selected incident — overview, timeline, runbooks, actions and review.
Four cards across the top summarize your reliability:
| Card | Meaning |
|---|---|
| Open Incidents | Incidents not yet resolved (open, investigating or mitigating). |
| Critical | Unresolved incidents at critical severity. |
| Avg MTTR | Mean time to resolve, over incidents resolved in the last 30 days. |
| Total Resolved | All incidents resolved to date. |
The incident list
Each row shows:
- Severity: critical (red), high (orange), medium (amber), low (blue).
- Status: open, investigating, mitigating, resolved.
- The title and affected service, cluster and namespace.
- Age, and time to resolve once resolved.
One incident, not a hundred alerts
When something breaks, it rarely breaks once. The incident manager groups events that share the same cluster, category and namespace within a five-minute window into a single incident. Ten pod crashes in three minutes in the payments namespace become one incident with ten source events — not ten pages to your on-call engineer.
Working an incident
Overview
Title, description, category (node_failure, pod_crash, security, performance), affected service, and the source events that opened it.
Timeline
Everything that happens to the incident is recorded:
| Event | Meaning |
|---|---|
| created | The incident was opened, by KubeOpera or by a person. |
| acknowledged | Someone took ownership. |
| runbook_started | An automated runbook began. |
| escalated | Severity was raised or the incident was escalated to another responder. |
| resolved | The incident was closed. |
Runbooks
Runbooks are automated response steps that run when an incident matches. Each step is one of kubectl, http, notify or wait, and shows its status (pending, executing, success, failed) and output as it runs. See Incident manager to write your own.
Actions
| Button | What it does |
|---|---|
| Acknowledge | Take ownership; notifications for this incident stop escalating. |
| Escalate | Raise severity or hand the incident to the next responder. |
| Investigate with AI | Launch an Incident Responder agent with this incident as context. |
| Resolve | Close the incident and record the time to resolve. |
Post-incident review
When an incident is resolved, select Generate review to create a post-incident review (PIR): a summary of what happened, the timeline, contributing factors, what was done, and follow-up actions. Edit it, then share it with your team. Reviews stay attached to the incident for future reference.
Open an incident yourself
Select New Incident, or use the API:
POST /api/incidents/api/v1/incidents
Content-Type: application/json
{
"cluster_id": "prod-us-east",
"title": "Elevated error rate on payments-api",
"description": "5xx rate above 2% for 10+ minutes",
"severity": "high",
"category": "performance",
"affected_service": "payments-api",
"namespace": "payments"
}
Next steps
- Incident manager — correlation, runbooks, notifications and API.
- Your first AI investigation — let an agent diagnose an incident.