Incident Manager
incident-manager provides the full incident lifecycle: correlation, runbook execution, post-incident reports, notifications, and MTTR tracking.
Domain Model
Incident
├── id, cluster_id, title, description
├── severity: low | medium | high | critical
├── status: open | investigating | mitigating | resolved
├── category: node_failure | pod_crash | security | performance
├── affected_service, namespace
├── source_event_ids[], runbook_id, assigned_to
├── acknowledged_at, mitigated_at, resolved_at
└── mttr_secs, post_mortem_url
IncidentEvent
├── incident_id, event_type, message
├── author: system | {user-email}
└── occurred_at
Runbook
├── id, cluster_id, name, category, description
├── auto_execute: bool
└── steps: []RunbookStep
RunbookStep
├── order, name
├── type: kubectl | http | notify | wait
├── command, url, method, payload
└── timeout_secs, continue_on_error
RunbookExecution
├── id, incident_id, runbook_id, started_by
├── status: pending | running | success | failed
├── step_results: []StepResult
├── started_at, finished_at
└── error
StepResult
├── order, name
├── status: pending | running | success | failed
├── output, error
└── started_at, finished_at
PostIncidentReport
├── id, incident_id
├── summary, root_cause, impact
├── action_items: []string
└── generated_at
Incident Correlation
A 5-minute sliding window is maintained in memory using sync.Map, keyed on {clusterID}:{category}:{namespace}. When a new event arrives:
- If a key exists and is within the 5-minute window → the event is appended to the existing incident as a new
IncidentEvent - If the key has expired or does not exist → a new incident is created
This prevents alert storms from generating hundreds of duplicate incidents for the same root cause.
Runbook Execution
Runbook steps execute sequentially in a background goroutine. The API call returns immediately with an execution ID — clients can poll /api/v1/executions/{execId} to track progress.
Step types:
| Type | What it does |
|---|---|
kubectl | Runs a kubectl command against the affected cluster |
http | Makes an HTTP request to an internal or external endpoint |
notify | Sends a notification to all configured channels |
wait | Pauses for timeout_secs seconds before proceeding |
If ContinueOnError: true, a failing step does not block subsequent steps. Each step produces a StepResult with full output and error text persisted to PostgreSQL.
Post-Incident Reports
POST /api/v1/incidents/{id}/pir triggers an AI-generated post-incident report:
- Loads the full incident timeline (events + runbook executions)
- Builds a chronological summary string
- Calls Claude Haiku (
claude-haiku-4-5-20251001) with the timeline - Parses the structured JSON response into
{summary, root_cause, impact, action_items[]} - Persists the
PostIncidentReportand returns it
The model is prompted to respond only with valid JSON, with a plain-text fallback if parsing fails. One report is stored per incident (ON CONFLICT (incident_id) DO UPDATE).
Notifications
All configured channels receive the same alert via a fan-out MultiNotifier. Each channel is activated independently by its environment variable — any combination is valid.
- Slack (
SLACK_WEBHOOK_URL) — posts a colour-coded attachment with severity, title, cluster, namespace, status, and a link to the KubeOpera incidents page - PagerDuty (
PAGERDUTY_INTEGRATION_KEY) — creates or resolves a PagerDuty incident via the Events API v2, using the incident ID as the dedup key. Severity is mapped:critical→critical,high→error,medium→warning,low→info. Set toresolveautomatically when the incident status changes toresolved. - Generic webhook (
WEBHOOK_URL) — POSTs a typed JSON payload to any HTTP endpoint. Supports an optionalWEBHOOK_SECRETsent as theX-Webhook-Secretheader. Theevent_typefield is one ofincident.created,incident.updated, orincident.resolved.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/incidents | List incidents (?cluster_id=&status=&limit=) |
POST | /api/v1/incidents | Create incident manually |
GET | /api/v1/incidents/{id} | Full incident detail + timeline |
POST | /api/v1/incidents/{id}/acknowledge | Mark as investigating |
POST | /api/v1/incidents/{id}/resolve | Close incident (records MTTR) |
GET | /api/v1/incidents/{id}/timeline | Ordered IncidentEvent list |
GET | /api/v1/incidents/metrics | MTTR avg, open count, critical count, total resolved |
GET | /api/v1/runbooks | List runbooks |
POST | /api/v1/runbooks | Create runbook |
PUT | /api/v1/runbooks/{id} | Update runbook |
DELETE | /api/v1/runbooks/{id} | Delete runbook |
POST | /api/v1/incidents/{id}/runbooks/{rbId}/execute | Execute runbook (async — returns execution ID) |
GET | /api/v1/incidents/{id}/executions | Runbook execution history for an incident |
GET | /api/v1/executions/{execId} | Execution detail with per-step results |
POST | /api/v1/incidents/{id}/pir | Generate AI post-incident report |
GET | /api/v1/incidents/{id}/pir | Retrieve generated post-incident report |
Environment Variables
| Variable | Description |
|---|---|
DATABASE_URL | PostgreSQL connection |
RABBITMQ_URL | RabbitMQ connection |
ANTHROPIC_API_KEY | Required for POST /pir (Claude Haiku post-incident reports) |
SLACK_WEBHOOK_URL | Optional — Slack incoming webhook URL |
PAGERDUTY_INTEGRATION_KEY | Optional — PagerDuty Events API v2 integration key |
WEBHOOK_URL | Optional — Generic HTTP webhook endpoint |
WEBHOOK_SECRET | Optional — Sent as X-Webhook-Secret header on webhook calls |
CORRELATION_WINDOW_SECS | Dedup window (default: 300) |
PORT | HTTP port (default: 8090) |