Anomaly Detector
Service: anomaly-detector · Port: 8088 · Database schema: anomaly
anomaly-detector gives you explicit, named rules for responding to anomalies. It receives anomaly events from k8s-monitor and the Optimizer, checks them against the alert rules you define, runs the self-healing action a rule specifies (automatically, or after approval), and records everything.
It complements the reactive AI pipeline: the pipeline adapts on its own; alert rules do exactly what you tell them for the workloads you care about most.
How it works
- An anomaly event arrives on
k8s.anomalies(from k8s-monitor or kubeopera-ai). - The detector stores it and finds matching alert rules for its cluster and metric.
- For each matching rule:
- if
auto_approveis true, it asks the action agent to run the rule's action; - if false, the action waits for approval in the action queue.
- if
- The result is recorded as a remediation and published to
k8s.selfheal, where the incident manager picks it up.
Alert rules
POST /api/v1/alert-rules
Content-Type: application/json
{
"cluster_id": "prod-us-east",
"name": "High CPU on API pods",
"metric": "cpu_usage_pct",
"condition": "z_score_gt",
"threshold": 2.5,
"severity": "high",
"namespace": "api",
"action": "scale_up",
"enabled": true,
"auto_approve": false
}
| Field | Values |
|---|---|
condition | gt (value above), lt (value below), z_score_gt (unusually high for this metric) |
action | notify, restart_pod, cordon_node, scale_up |
namespace, resource | Optional — limit the rule to specific workloads. |
auto_approve | Run the action automatically (true) or wait for approval (false). |
Default behaviour
With no custom rules, these defaults apply:
| When | Action |
|---|---|
| Pod restarts, severity ≥ high | Restart the pod. |
| Memory usage critical on a node | Cordon the node (requires approval). |
| CPU usage high on a deployment | Scale up by one replica. |
| Anything else | Notify only. |
caution
Cordoning stops new pods landing on a node but doesn't move existing ones. Use it for nodes under resource exhaustion, and drain the node if pods need to move.
Domain model
AnomalyEvent
├── id, cluster_id, source
├── metric, value, baseline, z_score
├── severity: low | medium | high | critical
├── namespace, resource, message
├── acknowledged, remediation_id
└── detected_at
AlertRule
├── id, cluster_id, name
├── metric, condition, threshold, severity
├── namespace?, resource?
├── action, enabled, auto_approve
Remediation
├── id, anomaly_id, rule_id, action_id
├── action_type, target_kind, target_name, namespace
├── status: pending | awaiting_approval | executing | success | failed
└── created_at, completed_at
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/anomalies | Anomaly events (?cluster_id=&severity=&start=&end=). |
POST | /api/v1/anomalies/{id}/acknowledge | Acknowledge an anomaly. |
GET | /api/v1/anomalies/stats | Counts by severity. |
GET · POST | /api/v1/alert-rules | List or create alert rules. |
PUT · DELETE | /api/v1/alert-rules/{id} | Update or delete a rule. |
GET | /api/v1/remediations | Remediation history. |
POST | /api/v1/remediations/{id}/retry | Retry a failed remediation. |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
ACTION_AGENT_SRV_BASE_URL | http://action-agent-srv:8094 | Where actions are executed. |
PORT | 8088 | HTTP port. |