Skip to main content
Version: 1.0

Incident Manager

incident-manager provides the full incident lifecycle: correlation, runbook execution, post-incident reports, notifications, and MTTR tracking.

Domain Model​

Incident
├── id, cluster_id, title, description
├── severity: low | medium | high | critical
├── status: open | investigating | mitigating | resolved
├── category: node_failure | pod_crash | security | performance
├── affected_service, namespace
├── source_event_ids[], runbook_id, assigned_to
├── acknowledged_at, mitigated_at, resolved_at
└── mttr_secs, post_mortem_url

IncidentEvent
├── incident_id, event_type, message
├── author: system | {user-email}
└── occurred_at

Runbook
├── id, cluster_id, name, category, description
├── auto_execute: bool
└── steps: []RunbookStep

RunbookStep
├── order, name
├── type: kubectl | http | notify | wait
├── command, url, method, payload
└── timeout_secs, continue_on_error

RunbookExecution
├── id, incident_id, runbook_id, started_by
├── status: pending | running | success | failed
├── step_results: []StepResult
├── started_at, finished_at
└── error

StepResult
├── order, name
├── status: pending | running | success | failed
├── output, error
└── started_at, finished_at

PostIncidentReport
├── id, incident_id
├── summary, root_cause, impact
├── action_items: []string
└── generated_at

Incident Correlation​

A 5-minute sliding window is maintained in memory using sync.Map, keyed on {clusterID}:{category}:{namespace}. When a new event arrives:

  • If a key exists and is within the 5-minute window → the event is appended to the existing incident as a new IncidentEvent
  • If the key has expired or does not exist → a new incident is created

This prevents alert storms from generating hundreds of duplicate incidents for the same root cause.

Runbook Execution​

Runbook steps execute sequentially in a background goroutine. The API call returns immediately with an execution ID — clients can poll /api/v1/executions/{execId} to track progress.

Step types:

TypeWhat it does
kubectlRuns a kubectl command against the affected cluster
httpMakes an HTTP request to an internal or external endpoint
notifySends a notification to all configured channels
waitPauses for timeout_secs seconds before proceeding

If ContinueOnError: true, a failing step does not block subsequent steps. Each step produces a StepResult with full output and error text persisted to PostgreSQL.

Post-Incident Reports​

POST /api/v1/incidents/{id}/pir triggers an AI-generated post-incident report:

  1. Loads the full incident timeline (events + runbook executions)
  2. Builds a chronological summary string
  3. Calls Claude Haiku (claude-haiku-4-5-20251001) with the timeline
  4. Parses the structured JSON response into {summary, root_cause, impact, action_items[]}
  5. Persists the PostIncidentReport and returns it

The model is prompted to respond only with valid JSON, with a plain-text fallback if parsing fails. One report is stored per incident (ON CONFLICT (incident_id) DO UPDATE).

Notifications​

All configured channels receive the same alert via a fan-out MultiNotifier. Each channel is activated independently by its environment variable — any combination is valid.

  • Slack (SLACK_WEBHOOK_URL) — posts a colour-coded attachment with severity, title, cluster, namespace, status, and a link to the KubeOpera incidents page
  • PagerDuty (PAGERDUTY_INTEGRATION_KEY) — creates or resolves a PagerDuty incident via the Events API v2, using the incident ID as the dedup key. Severity is mapped: critical→critical, high→error, medium→warning, low→info. Set to resolve automatically when the incident status changes to resolved.
  • Generic webhook (WEBHOOK_URL) — POSTs a typed JSON payload to any HTTP endpoint. Supports an optional WEBHOOK_SECRET sent as the X-Webhook-Secret header. The event_type field is one of incident.created, incident.updated, or incident.resolved.

REST API​

MethodPathDescription
GET/api/v1/incidentsList incidents (?cluster_id=&status=&limit=)
POST/api/v1/incidentsCreate incident manually
GET/api/v1/incidents/{id}Full incident detail + timeline
POST/api/v1/incidents/{id}/acknowledgeMark as investigating
POST/api/v1/incidents/{id}/resolveClose incident (records MTTR)
GET/api/v1/incidents/{id}/timelineOrdered IncidentEvent list
GET/api/v1/incidents/metricsMTTR avg, open count, critical count, total resolved
GET/api/v1/runbooksList runbooks
POST/api/v1/runbooksCreate runbook
PUT/api/v1/runbooks/{id}Update runbook
DELETE/api/v1/runbooks/{id}Delete runbook
POST/api/v1/incidents/{id}/runbooks/{rbId}/executeExecute runbook (async — returns execution ID)
GET/api/v1/incidents/{id}/executionsRunbook execution history for an incident
GET/api/v1/executions/{execId}Execution detail with per-step results
POST/api/v1/incidents/{id}/pirGenerate AI post-incident report
GET/api/v1/incidents/{id}/pirRetrieve generated post-incident report

Environment Variables​

VariableDescription
DATABASE_URLPostgreSQL connection
RABBITMQ_URLRabbitMQ connection
ANTHROPIC_API_KEYRequired for POST /pir (Claude Haiku post-incident reports)
SLACK_WEBHOOK_URLOptional — Slack incoming webhook URL
PAGERDUTY_INTEGRATION_KEYOptional — PagerDuty Events API v2 integration key
WEBHOOK_URLOptional — Generic HTTP webhook endpoint
WEBHOOK_SECRETOptional — Sent as X-Webhook-Secret header on webhook calls
CORRELATION_WINDOW_SECSDedup window (default: 300)
PORTHTTP port (default: 8090)