Analysis Agent
The analysis agent is the statistical brain of the reactive pipeline. It consumes telemetry snapshots and determines whether any metrics indicate abnormal behaviour using Z-score anomaly detection.
Z-Score Detection
For each metric in every incoming TelemetrySnapshot, the analysis agent maintains a rolling window of the last N values (default: 60 points). When a new value arrives:
z = (current_value - rolling_mean) / rolling_std_dev
If Z exceeds the threshold, an Anomaly is recorded with the observed value, baseline (mean), Z-score, and severity.
Default thresholds (per cluster, adaptive):
| Z-score | Severity |
|---|---|
| ≥ 2.5 | high |
| ≥ 3.5 | critical |
| < 2.5 | below threshold (no anomaly) |
Thresholds are stored per-cluster-per-metric in memory and adjusted by the feedback loop.
AnalysisResult
Each telemetry message produces one AnalysisResult:
type AnalysisResult struct {
ID string
ClusterID string
RiskScore float64 // 0–100, weighted sum of anomaly severities
Anomalies []Anomaly
Decisions []Decision // what actions to take
Summary string
CreatedAt time.Time
}
RiskScore formula:
risk_score = min(100, Σ (
critical_count × 25 +
high_count × 10 +
medium_count × 5 +
low_count × 1
))
Decisions
For each anomaly, the analysis agent selects a Decision.Type (its own short internal names — scale/restart/cordon/notify — distinct from the longer scale_deployment/restart_pod/cordon_node action-type names action-agent-srv uses downstream) based on the metric name and severity, and sets AutoApprove per-decision, not uniformly:
| Metric contains | Condition | Decision | Auto-approved? |
|---|---|---|---|
cpu_usage | severity ≥ high | scale | Only when severity is critical |
crash_loop | severity ≥ high | restart | Always |
memory_usage | severity = critical | cordon | Never — always needs human approval |
| anything else | any severity | notify | Always (trivially — a notify has no destructive action to gate) |
Published Messages
analysis.decisionsexchange — aDecisionPublishMessagecontaining decisions for the action agent, published on every telemetry cycle.analysis.insightsexchange — anInsightPublishMessagecontaining risk score, summary, and anomaly list for the recommendation agent and agent-runtime. This one is gated, not published unconditionally: it only goes out when at least one anomaly was found and the risk score meets a minimum (INSIGHT_PUBLISH_MIN_RISK_SCORE, default 5 — the lowest non-zero value the risk-score formula above can produce). Before this gate existed, an insight went out every ~30 seconds per cluster regardless of whether anything notable had happened, which meantrecommendation-agent-srvandagent-runtime's auto-trigger were both doing real work on essentially empty signal most of the time.
Feedback Signal Consumption — built, but not wired up
The threshold-adjustment logic is real and tested:
// reinforcement (action improved the metric)
threshold[clusterID][metric] -= 0.05
// correction (action made the metric worse)
threshold[clusterID][metric] += 0.10
// Bounds: never below 1.5 or above 5.0
But analysis-agent-srv never actually starts a consumer for the feedback.signals exchange — this function is only reachable via an internal HTTP endpoint that nothing else in the platform calls. feedback-agent-srv genuinely publishes real FeedbackSignal messages to that exchange (see Feedback Agent), but as of today, nothing on the receiving end ever reads them — so the adaptive-per-cluster-per-metric threshold loop this section describes has never actually run in practice. Both halves of the loop are real, working code; only the RabbitMQ wiring between them is missing.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/results | Recent analysis results (?cluster_id=&limit=) |
GET | /api/v1/results/latest | Most recent result for a cluster |
GET | /healthz | Health check |
Environment Variables
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection |
RABBITMQ_URL | — | RabbitMQ connection |
ANOMALY_WINDOW_SIZE | 60 | Rolling window size |
ANOMALY_Z_THRESHOLD_HIGH | 2.5 | Z-score for high severity |
ANOMALY_Z_THRESHOLD_CRITICAL | 3.5 | Z-score for critical severity |
INSIGHT_PUBLISH_MIN_RISK_SCORE | 5 | Minimum risk score required before an analysis.insights message is published |
PORT | 8093 | HTTP port |