Incident Manager
Service: incident-manager · Port: 8090 · Database schema: incidents
incident-manager handles incidents from the first signal to the post-incident review. It turns streams of anomalies, self-healing events, security findings and failed deploys into a manageable set of incidents, runs runbooks, notifies the right people, and tracks how quickly you recover.
For the user's view, see Incidents.
Where incidents come from
incident-manager subscribes to:
| Exchange | Opens incidents for… |
|---|---|
k8s.anomalies | High and critical anomalies. |
k8s.selfheal | Failed self-healing actions (and records successful ones on existing incidents). |
security.posture | Critical security findings. |
cicd.events | Failed deploys to production, and links recent deploys to open incidents. |
nodes.events | Node failures and interruptions. |
People and agents can also open incidents through the API.
Correlation
To stop alert storms, related signals become one incident. Events are grouped by {cluster}:{category}:{namespace} within a sliding window (five minutes by default):
- if an open incident matches, the event is added to its timeline;
- otherwise a new incident is opened.
The correlation state is shared across replicas, so grouping works the same however many instances run.
Runbooks
A runbook is a sequence of steps that runs automatically when an incident matches (auto_execute: true) or on demand.
{
"name": "Restart crash-looping payments pods",
"category": "pod_crash",
"auto_execute": true,
"steps": [
{ "order": 1, "name": "Capture logs", "type": "kubectl",
"command": "logs -n {{namespace}} -l app={{affected_service}} --previous --tail=200" },
{ "order": 2, "name": "Restart", "type": "kubectl",
"command": "rollout restart deployment/{{affected_service}} -n {{namespace}}" },
{ "order": 3, "name": "Wait", "type": "wait", "timeout_secs": 120 },
{ "order": 4, "name": "Tell the team", "type": "notify" }
]
}
| Step type | What it does |
|---|---|
kubectl | Runs a kubectl command against the incident's cluster. |
http | Calls an HTTP endpoint (with url, method, payload). |
notify | Sends a message to the configured channels. |
wait | Pauses for timeout_secs. |
Steps run in order; set continue_on_error: true to keep going after a failure. Execution is asynchronous — the API returns an execution ID, and each step's status and output are recorded. Placeholders such as {{namespace}} are filled in from the incident.
Notifications
Every channel you configure receives each alert:
- Slack — a color-coded message with severity, title, cluster, namespace, status and a link.
- PagerDuty — incidents are created and resolved through Events API v2, using the incident ID for deduplication; severities map
critical→critical,high→error,medium→warning,low→info. - Microsoft Teams and email.
- Webhooks — through KubeOpera's webhook subscriptions (
incident.*events).
Escalation policies re-notify, or notify the next responder, when an incident isn't acknowledged in time.
Post-incident reviews
POST /api/v1/incidents/{id}/pir drafts a review with Claude from the incident's full timeline — events, runbook runs, linked deploys and agent investigations:
{
"summary": "…",
"root_cause": "…",
"impact": "…",
"timeline": [ … ],
"action_items": [ "…" ]
}
The review uses the tenant's AI credential, is editable in the dashboard, and stays attached to the incident. Generating again updates it.
Reliability metrics
| Metric | Meaning |
|---|---|
| MTTA | Mean time to acknowledge. |
| MTTR | Mean time to resolve. |
| MTBF | Mean time between failures, per service. |
| SLA compliance | Share of incidents resolved within their severity's target. |
Domain model
Incident
├── id, cluster_id, tenant_id, title, description
├── severity: low | medium | high | critical
├── status: open | investigating | mitigating | resolved
├── category: node_failure | pod_crash | security | performance | deployment
├── affected_service, namespace
├── source_event_ids[], linked_deploys[], agent_run_ids[]
├── assigned_to, acknowledged_at, mitigated_at, resolved_at
└── mttr_secs
Runbook → steps: []{ order, name, type, command | url, timeout_secs, continue_on_error }
RunbookExecution → status, step_results[], started_by, started_at, finished_at
PostIncidentReport → summary, root_cause, impact, timeline, action_items[]
REST API
| Method | Path | Description |
|---|---|---|
GET · POST | /api/v1/incidents | List (?cluster_id=&status=&limit=) or open an incident. |
GET | /api/v1/incidents/{id} | Detail and timeline. |
POST | /api/v1/incidents/{id}/acknowledge | Acknowledge. |
POST | /api/v1/incidents/{id}/escalate | Escalate ({ "severity"?, "assign_to"? }). |
POST | /api/v1/incidents/{id}/resolve | Resolve. |
GET | /api/v1/incidents/{id}/timeline | Timeline events. |
GET | /api/v1/incidents/metrics | MTTA, MTTR, open and critical counts, resolved total. |
GET · POST | /api/v1/runbooks | List or create runbooks. |
PUT · DELETE | /api/v1/runbooks/{id} | Update or delete a runbook. |
POST | /api/v1/incidents/{id}/runbooks/{rbId}/execute | Run a runbook. |
GET | /api/v1/incidents/{id}/executions | Runbook runs for an incident. |
GET | /api/v1/executions/{execId} | Execution detail with step results. |
POST · GET | /api/v1/incidents/{id}/pir | Generate or read the post-incident review. |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
CORRELATION_WINDOW_SECS | 300 | Correlation window. |
SLACK_WEBHOOK_URL | — | Slack channel. |
PAGERDUTY_INTEGRATION_KEY | — | PagerDuty Events API v2 key. |
TEAMS_WEBHOOK_URL | — | Microsoft Teams channel. |
SMTP_URL / NOTIFY_EMAIL_TO | — | Email notifications. |
AUTH_SERVICE_BASE_URL / AI_CREDENTIAL_INTERNAL_API_KEY | — | AI credential resolution for reviews. |
PORT | 8090 | HTTP port. |