Reactive AI Pipeline
The reactive pipeline detects problems, takes action, and measures results — detection, analysis, and action form a real, working chain, running automatically with no human intervention needed for auto-approved decisions. The one piece that isn't real yet is the "learns" part its name implies: the final loop, where feedback about an action's outcome tunes the system's own detection thresholds, has the logic built on both ends but isn't wired together — see below for exactly where.
Event Flow
Reactive AI Pipeline — Event Flow
Autonomous detection → action → feedback. The loop back to detection isn't wired up yet.
Adaptive Thresholds — designed, not yet running
The analysis agent maintains per-cluster, per-metric Z-score thresholds in memory, and has real, tested logic to adjust them:
| Feedback type | Meaning | Threshold change |
|---|---|---|
reinforcement | The triggering action executed successfully | -0.05 (more sensitive) |
correction | The triggering action failed or errored | +0.10 (less sensitive) |
If this loop were actually running, a cluster whose auto-approved actions consistently succeed would gradually receive faster alerts (lower threshold), while one where actions keep failing would self-quieten. In practice, analysis-agent-srv never starts a consumer for the feedback.signals exchange this table's inputs are published to, so this adjustment code has never executed outside of tests. Worth being precise about what "the action improved the metric" means, too: feedback-agent-srv derives reinforcement/correction purely from whether the Kubernetes action itself succeeded or failed — there's no follow-up telemetry comparison checking whether cluster health actually got better. See Analysis Agent and Feedback Agent for the full detail on both halves.
Telemetry Message Schema
type TelemetryPublishMessage struct {
SnapshotID string
ClusterID string
Timestamp time.Time
Snapshot TelemetrySnapshot
}
type TelemetrySnapshot struct {
ClusterID string
Source string // "k8s-monitor" | "security-api" | "cicd-gateway"
HealthScore float64
CPUUsagePct float64
MemoryUsagePct float64
ReadyNodeCount int
NodeCount int
PodCount int
FailedPodCount int
CrashLoopCount int
APIServerLatencyMs float64
NetworkRxBytesPS float64
NetworkTxBytesPS float64
SecurityPostureScore float64 // host-cluster targets only
}
See Observability Agent for how ClusterID gets populated — it's a real cluster ID for host-cluster targets, but a tenant's vCluster namespace for the per-tenant targets this service also discovers and collects from.
Decision Types
The analysis agent produces decisions of the following types — see Analysis Agent for the exact metric/severity conditions and which decisions are auto-approved by default:
| Type | Default action (via action-agent-srv) |
|---|---|
scale | Scale the target Deployment's replicas |
restart | Delete the target Pod |
cordon | Patch the target Node unschedulable |
notify | Log only; no Kubernetes action taken |
Configuring Alert Rules
This section covers a separate, related mechanism — anomaly-detector's own AlertRule system, not the analysis-agent Z-score pipeline described above. anomaly-detector consumes its own anomaly events from the k8s.anomalies exchange (a different exchange from analysis-agent's analysis.decisions/analysis.insights) and evaluates them against configured rules:
POST /api/anomalies/api/v1/alert-rules
Content-Type: application/json
{
"cluster_id": "prod-us-east",
"name": "High CPU on API pods",
"metric": "cpu_usage_pct",
"condition": "z_score_gt",
"threshold": 2.5,
"severity": "high",
"action": "scale_up",
"enabled": true,
"auto_approve": false
}
Set auto_approve: true to enable fully automated remediation. When false, the anomaly is logged and visible in the UI but no Kubernetes action is taken.