Anomaly Detector
anomaly-detector consumes anomaly events published by k8s-monitor, evaluates alert rules, optionally executes Kubernetes self-healing actions, and publishes outcomes to incident-manager.
Domain Model
AnomalyEvent
├── id, cluster_id
├── metric: cpu_usage_pct | memory_usage_pct | crash_loop_count | ...
├── value, baseline, z_score
├── severity: low | medium | high | critical
├── namespace, resource, message
├── acknowledged, healing_id
└── detected_at
AlertRule
├── id, cluster_id, name
├── metric, condition: gt | lt | z_score_gt
├── threshold, severity
├── action: notify | restart_pod | cordon_node | scale_up
└── enabled, auto_approve
RemediationLog
├── id, anomaly_id, cluster_id
├── action_type, target_kind, target_name, namespace
├── status: pending | executing | success | failed
└── executed_at, completed_at
Self-Healing Decision Tree
When an anomaly event arrives and a matching AlertRule has AutoApprove: true:
| Condition | Action | Kubernetes call |
|---|---|---|
pod_restarts + severity ≥ high | restart_pod | Delete pod (controller recreates) |
memory_usage + critical + node | cordon_node | Patch spec.unschedulable: true |
cpu_usage + high + deployment | scale_deployment | Increment replicas by 1 |
| default | notify | No K8s mutation; log only |
Actions are executed by the action-agent-srv, not inline. The anomaly-detector publishes a DecisionPublishMessage which the action agent picks up.
caution
cordon_node prevents new pods from being scheduled on the affected node but does not evict existing pods. Use this action only when the node is experiencing resource exhaustion that could affect pod stability.
Kubernetes RBAC Requirements
The anomaly-detector service account needs:
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["list", "get", "delete"]
- apiGroups: [""]
resources: ["nodes"]
verbs: ["list", "get", "patch"]
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["list", "get", "patch", "update"]
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/anomalies | List events (?cluster_id=&severity=&start=&end=) |
POST | /api/v1/anomalies/{id}/acknowledge | Mark event reviewed |
GET | /api/v1/anomalies/stats | Severity counts for dashboard widget |
GET | /api/v1/alert-rules | List alert rules |
POST | /api/v1/alert-rules | Create alert rule |
PUT | /api/v1/alert-rules/{id} | Update alert rule |
DELETE | /api/v1/alert-rules/{id} | Delete alert rule |
GET | /api/v1/remediations | Self-healing action log |
POST | /api/v1/remediations/{id}/retry | Retry a failed remediation |
Environment Variables
| Variable | Description |
|---|---|
DATABASE_URL | PostgreSQL connection |
RABBITMQ_URL | RabbitMQ connection |
KUBECONFIG | Path to kubeconfig (or in-cluster) |
PORT | HTTP port (default: 8088) |