Skip to main content
Version: 2.0

Incident Manager

Service: incident-manager · Port: 8090 · Database schema: incidents

incident-manager handles incidents from the first signal to the post-incident review. It turns streams of anomalies, self-healing events, security findings and failed deploys into a manageable set of incidents, runs runbooks, notifies the right people, and tracks how quickly you recover.

For the user's view, see Incidents.

Where incidents come from​

incident-manager subscribes to:

ExchangeOpens incidents for…
k8s.anomaliesHigh and critical anomalies.
k8s.selfhealFailed self-healing actions (and records successful ones on existing incidents).
security.postureCritical security findings.
cicd.eventsFailed deploys to production, and links recent deploys to open incidents.
nodes.eventsNode failures and interruptions.

People and agents can also open incidents through the API.

Correlation​

To stop alert storms, related signals become one incident. Events are grouped by {cluster}:{category}:{namespace} within a sliding window (five minutes by default):

  • if an open incident matches, the event is added to its timeline;
  • otherwise a new incident is opened.

The correlation state is shared across replicas, so grouping works the same however many instances run.

Runbooks​

A runbook is a sequence of steps that runs automatically when an incident matches (auto_execute: true) or on demand.

{
"name": "Restart crash-looping payments pods",
"category": "pod_crash",
"auto_execute": true,
"steps": [
{ "order": 1, "name": "Capture logs", "type": "kubectl",
"command": "logs -n {{namespace}} -l app={{affected_service}} --previous --tail=200" },
{ "order": 2, "name": "Restart", "type": "kubectl",
"command": "rollout restart deployment/{{affected_service}} -n {{namespace}}" },
{ "order": 3, "name": "Wait", "type": "wait", "timeout_secs": 120 },
{ "order": 4, "name": "Tell the team", "type": "notify" }
]
}
Step typeWhat it does
kubectlRuns a kubectl command against the incident's cluster.
httpCalls an HTTP endpoint (with url, method, payload).
notifySends a message to the configured channels.
waitPauses for timeout_secs.

Steps run in order; set continue_on_error: true to keep going after a failure. Execution is asynchronous — the API returns an execution ID, and each step's status and output are recorded. Placeholders such as {{namespace}} are filled in from the incident.

Notifications​

Every channel you configure receives each alert:

  • Slack — a color-coded message with severity, title, cluster, namespace, status and a link.
  • PagerDuty — incidents are created and resolved through Events API v2, using the incident ID for deduplication; severities map critical→critical, high→error, medium→warning, low→info.
  • Microsoft Teams and email.
  • Webhooks — through KubeOpera's webhook subscriptions (incident.* events).

Escalation policies re-notify, or notify the next responder, when an incident isn't acknowledged in time.

Post-incident reviews​

POST /api/v1/incidents/{id}/pir drafts a review with Claude from the incident's full timeline — events, runbook runs, linked deploys and agent investigations:

{
"summary": "…",
"root_cause": "…",
"impact": "…",
"timeline": [ … ],
"action_items": [ "…" ]
}

The review uses the tenant's AI credential, is editable in the dashboard, and stays attached to the incident. Generating again updates it.

Reliability metrics​

MetricMeaning
MTTAMean time to acknowledge.
MTTRMean time to resolve.
MTBFMean time between failures, per service.
SLA complianceShare of incidents resolved within their severity's target.

Domain model​

Incident
├── id, cluster_id, tenant_id, title, description
├── severity: low | medium | high | critical
├── status: open | investigating | mitigating | resolved
├── category: node_failure | pod_crash | security | performance | deployment
├── affected_service, namespace
├── source_event_ids[], linked_deploys[], agent_run_ids[]
├── assigned_to, acknowledged_at, mitigated_at, resolved_at
└── mttr_secs

Runbook → steps: []{ order, name, type, command | url, timeout_secs, continue_on_error }
RunbookExecution → status, step_results[], started_by, started_at, finished_at
PostIncidentReport → summary, root_cause, impact, timeline, action_items[]

REST API​

MethodPathDescription
GET · POST/api/v1/incidentsList (?cluster_id=&status=&limit=) or open an incident.
GET/api/v1/incidents/{id}Detail and timeline.
POST/api/v1/incidents/{id}/acknowledgeAcknowledge.
POST/api/v1/incidents/{id}/escalateEscalate ({ "severity"?, "assign_to"? }).
POST/api/v1/incidents/{id}/resolveResolve.
GET/api/v1/incidents/{id}/timelineTimeline events.
GET/api/v1/incidents/metricsMTTA, MTTR, open and critical counts, resolved total.
GET · POST/api/v1/runbooksList or create runbooks.
PUT · DELETE/api/v1/runbooks/{id}Update or delete a runbook.
POST/api/v1/incidents/{id}/runbooks/{rbId}/executeRun a runbook.
GET/api/v1/incidents/{id}/executionsRunbook runs for an incident.
GET/api/v1/executions/{execId}Execution detail with step results.
POST · GET/api/v1/incidents/{id}/pirGenerate or read the post-incident review.
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—RabbitMQ connection.
CORRELATION_WINDOW_SECS300Correlation window.
SLACK_WEBHOOK_URL—Slack channel.
PAGERDUTY_INTEGRATION_KEY—PagerDuty Events API v2 key.
TEAMS_WEBHOOK_URL—Microsoft Teams channel.
SMTP_URL / NOTIFY_EMAIL_TO—Email notifications.
AUTH_SERVICE_BASE_URL / AI_CREDENTIAL_INTERNAL_API_KEY—AI credential resolution for reviews.
PORT8090HTTP port.