Skip to main content
Version: 2.0

Analysis Agent

Service: analysis-agent-srv · Port: 8093

The analysis agent is the decision-maker of the reactive pipeline. For every telemetry snapshot it answers three questions:

  1. Is anything abnormal? — anomaly detection with adaptive Z-scores.
  2. How worried should we be? — a risk score from 0 to 100.
  3. What should we do about it? — decisions for the action agent, and insights for the recommendation agent and agent runtime.

It then learns from the outcome of its decisions, so each cluster's definition of "abnormal" improves over time.

Detecting anomalies​

For each metric of each target, the agent keeps a rolling window of recent values (60 by default — about 30 minutes at the default collection interval). When a new value arrives it computes how unusual it is:

z = (current_value − rolling_mean) / rolling_std_dev

A large z means the value is far from normal for this metric on this cluster. If z passes the threshold, the agent records an anomaly with the observed value, the baseline (mean), the Z-score and a severity.

Starting thresholds (each cluster's thresholds then adapt — see below):

Z-scoreSeverity
≥ 3.5critical
≥ 2.5high
< 2.5normal — no anomaly
Why Z-scores?

A fixed rule like "CPU above 80%" is wrong for half your workloads. A Z-score asks "is this unusual for you?" — so a batch cluster that always runs hot isn't flagged, while a quiet API that suddenly doubles its CPU is.

Scoring risk​

Each snapshot produces one AnalysisResult:

type AnalysisResult struct {
ID string
ClusterID string
RiskScore float64 // 0–100
Anomalies []Anomaly
Decisions []Decision
Summary string
CreatedAt time.Time
}

The risk score weights anomalies by severity:

risk_score = min(100,
25 × critical_count +
10 × high_count +
5 × medium_count +
1 × low_count)

A single critical anomaly scores 25; three critical anomalies together push the score past 70, which launches an automatic SRE investigation.

Decisions​

For each anomaly the agent chooses a decision based on the metric and severity, and whether it may run automatically (AutoApprove). The defaults are:

MetricConditionDecisionRuns automatically?
cpu_usageseverity ≥ highscaleWhen critical; otherwise waits for approval
crash_loopseverity ≥ highrestartYes
memory_usageseverity = criticalcordonNo — always waits for approval
anything elseanynotifyYes (no cluster change is made)

Decisions that don't run automatically appear in the approval queue on /agents/actions. You can change which decisions run automatically, per cluster, with decision rules:

PUT /api/v1/decision-rules/prod-us-east
Content-Type: application/json

{
"rules": [
{ "metric": "cpu_usage", "min_severity": "high", "decision": "scale", "auto_approve": true },
{ "metric": "memory_usage", "min_severity": "critical", "decision": "cordon", "auto_approve": false }
]
}

Publishing results​

ExchangeMessageWhen
analysis.decisionsDecisionPublishMessage — the decisions for the action agent.Every cycle with at least one decision.
analysis.insightsInsightPublishMessage — risk score, summary and anomalies.When there is at least one anomaly and the risk score reaches INSIGHT_PUBLISH_MIN_RISK_SCORE (default 5).

The insight threshold keeps downstream AI work focused on signal: the recommendation agent and agent runtime only do work when something notable has happened.

Learning from feedback​

The analysis agent consumes FeedbackSignal messages from feedback.signals and adjusts the threshold for the signal's cluster and metric:

// reinforcement — the action improved the metric
threshold[clusterID][metric] -= 0.05 * magnitude

// correction — the action failed or didn't improve the metric
threshold[clusterID][metric] += 0.10 * magnitude

// thresholds are kept between 1.5 and 5.0

magnitude (0–1) reflects how strongly the metric responded, so a decisive result moves the threshold more than a marginal one. Thresholds are persisted, so they survive restarts, and you can see each cluster's current thresholds on /agents/feedback.

The effect: a cluster whose automatic actions consistently help gets faster alerts; a cluster where actions keep failing to help becomes quieter — instead of repeating an action that doesn't work.

REST API​

MethodPathDescription
GET/api/v1/resultsRecent analysis results (?cluster_id=&limit=).
GET/api/v1/results/latestThe most recent result for a cluster.
GET/api/v1/thresholdsCurrent adaptive thresholds (?cluster_id=).
GET / PUT/api/v1/decision-rules/{cluster_id}Read or change a cluster's decision rules.
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
ANOMALY_WINDOW_SIZE60Rolling window size, in data points.
ANOMALY_Z_THRESHOLD_HIGH2.5Starting Z-score for high severity.
ANOMALY_Z_THRESHOLD_CRITICAL3.5Starting Z-score for critical severity.
THRESHOLD_MIN / THRESHOLD_MAX1.5 / 5.0Bounds for adaptive thresholds.
INSIGHT_PUBLISH_MIN_RISK_SCORE5Minimum risk score to publish an insight.
DATABASE_URL—PostgreSQL connection.
RABBITMQ_URL—RabbitMQ connection.
PORT8093HTTP port.