Skip to main content
Version: 1.0

Analysis Agent

The analysis agent is the statistical brain of the reactive pipeline. It consumes telemetry snapshots and determines whether any metrics indicate abnormal behaviour using Z-score anomaly detection.

Z-Score Detection​

For each metric in every incoming TelemetrySnapshot, the analysis agent maintains a rolling window of the last N values (default: 60 points). When a new value arrives:

z = (current_value - rolling_mean) / rolling_std_dev

If Z exceeds the threshold, an Anomaly is recorded with the observed value, baseline (mean), Z-score, and severity.

Default thresholds (per cluster, adaptive):

Z-scoreSeverity
≥ 2.5high
≥ 3.5critical
< 2.5below threshold (no anomaly)

Thresholds are stored per-cluster-per-metric in memory and adjusted by the feedback loop.

AnalysisResult​

Each telemetry message produces one AnalysisResult:

type AnalysisResult struct {
ID string
ClusterID string
RiskScore float64 // 0–100, weighted sum of anomaly severities
Anomalies []Anomaly
Decisions []Decision // what actions to take
Summary string
CreatedAt time.Time
}

RiskScore formula:

risk_score = min(100, Σ (
critical_count × 25 +
high_count × 10 +
medium_count × 5 +
low_count × 1
))

Decisions​

For each anomaly, the analysis agent selects a Decision.Type (its own short internal names — scale/restart/cordon/notify — distinct from the longer scale_deployment/restart_pod/cordon_node action-type names action-agent-srv uses downstream) based on the metric name and severity, and sets AutoApprove per-decision, not uniformly:

Metric containsConditionDecisionAuto-approved?
cpu_usageseverity ≥ highscaleOnly when severity is critical
crash_loopseverity ≥ highrestartAlways
memory_usageseverity = criticalcordonNever — always needs human approval
anything elseany severitynotifyAlways (trivially — a notify has no destructive action to gate)

Published Messages​

  • analysis.decisions exchange — a DecisionPublishMessage containing decisions for the action agent, published on every telemetry cycle.
  • analysis.insights exchange — an InsightPublishMessage containing risk score, summary, and anomaly list for the recommendation agent and agent-runtime. This one is gated, not published unconditionally: it only goes out when at least one anomaly was found and the risk score meets a minimum (INSIGHT_PUBLISH_MIN_RISK_SCORE, default 5 — the lowest non-zero value the risk-score formula above can produce). Before this gate existed, an insight went out every ~30 seconds per cluster regardless of whether anything notable had happened, which meant recommendation-agent-srv and agent-runtime's auto-trigger were both doing real work on essentially empty signal most of the time.

Feedback Signal Consumption — built, but not wired up​

The threshold-adjustment logic is real and tested:

// reinforcement (action improved the metric)
threshold[clusterID][metric] -= 0.05

// correction (action made the metric worse)
threshold[clusterID][metric] += 0.10

// Bounds: never below 1.5 or above 5.0

But analysis-agent-srv never actually starts a consumer for the feedback.signals exchange — this function is only reachable via an internal HTTP endpoint that nothing else in the platform calls. feedback-agent-srv genuinely publishes real FeedbackSignal messages to that exchange (see Feedback Agent), but as of today, nothing on the receiving end ever reads them — so the adaptive-per-cluster-per-metric threshold loop this section describes has never actually run in practice. Both halves of the loop are real, working code; only the RabbitMQ wiring between them is missing.

REST API​

MethodPathDescription
GET/api/v1/resultsRecent analysis results (?cluster_id=&limit=)
GET/api/v1/results/latestMost recent result for a cluster
GET/healthzHealth check

Environment Variables​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection
RABBITMQ_URL—RabbitMQ connection
ANOMALY_WINDOW_SIZE60Rolling window size
ANOMALY_Z_THRESHOLD_HIGH2.5Z-score for high severity
ANOMALY_Z_THRESHOLD_CRITICAL3.5Z-score for critical severity
INSIGHT_PUBLISH_MIN_RISK_SCORE5Minimum risk score required before an analysis.insights message is published
PORT8093HTTP port