Skip to main content
Version: 2.0

Anomaly Detector

Service: anomaly-detector · Port: 8088 · Database schema: anomaly

anomaly-detector gives you explicit, named rules for responding to anomalies. It receives anomaly events from k8s-monitor and the Optimizer, checks them against the alert rules you define, runs the self-healing action a rule specifies (automatically, or after approval), and records everything.

It complements the reactive AI pipeline: the pipeline adapts on its own; alert rules do exactly what you tell them for the workloads you care about most.

How it works​

  1. An anomaly event arrives on k8s.anomalies (from k8s-monitor or kubeopera-ai).
  2. The detector stores it and finds matching alert rules for its cluster and metric.
  3. For each matching rule:
    • if auto_approve is true, it asks the action agent to run the rule's action;
    • if false, the action waits for approval in the action queue.
  4. The result is recorded as a remediation and published to k8s.selfheal, where the incident manager picks it up.

Alert rules​

POST /api/v1/alert-rules
Content-Type: application/json

{
"cluster_id": "prod-us-east",
"name": "High CPU on API pods",
"metric": "cpu_usage_pct",
"condition": "z_score_gt",
"threshold": 2.5,
"severity": "high",
"namespace": "api",
"action": "scale_up",
"enabled": true,
"auto_approve": false
}
FieldValues
conditiongt (value above), lt (value below), z_score_gt (unusually high for this metric)
actionnotify, restart_pod, cordon_node, scale_up
namespace, resourceOptional — limit the rule to specific workloads.
auto_approveRun the action automatically (true) or wait for approval (false).

Default behaviour​

With no custom rules, these defaults apply:

WhenAction
Pod restarts, severity ≥ highRestart the pod.
Memory usage critical on a nodeCordon the node (requires approval).
CPU usage high on a deploymentScale up by one replica.
Anything elseNotify only.
caution

Cordoning stops new pods landing on a node but doesn't move existing ones. Use it for nodes under resource exhaustion, and drain the node if pods need to move.

Domain model​

AnomalyEvent
├── id, cluster_id, source
├── metric, value, baseline, z_score
├── severity: low | medium | high | critical
├── namespace, resource, message
├── acknowledged, remediation_id
└── detected_at

AlertRule
├── id, cluster_id, name
├── metric, condition, threshold, severity
├── namespace?, resource?
├── action, enabled, auto_approve

Remediation
├── id, anomaly_id, rule_id, action_id
├── action_type, target_kind, target_name, namespace
├── status: pending | awaiting_approval | executing | success | failed
└── created_at, completed_at

REST API​

MethodPathDescription
GET/api/v1/anomaliesAnomaly events (?cluster_id=&severity=&start=&end=).
POST/api/v1/anomalies/{id}/acknowledgeAcknowledge an anomaly.
GET/api/v1/anomalies/statsCounts by severity.
GET · POST/api/v1/alert-rulesList or create alert rules.
PUT · DELETE/api/v1/alert-rules/{id}Update or delete a rule.
GET/api/v1/remediationsRemediation history.
POST/api/v1/remediations/{id}/retryRetry a failed remediation.
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—RabbitMQ connection.
ACTION_AGENT_SRV_BASE_URLhttp://action-agent-srv:8094Where actions are executed.
PORT8088HTTP port.