Reactive AI Pipeline
The reactive pipeline is KubeOpera's always-on first responder. Every 30 seconds it looks at every cluster, decides whether anything is wrong, fixes what it's allowed to fix, checks whether the fix worked, and uses that result to get better at spotting real problems.
It does all of this without a person in the loop for actions you've approved in advance — and every step is recorded so you can see exactly what it did and why.
The loop
Reactive AI Pipeline — Event Flow
Autonomous detection → action → feedback, with feedback tuning detection in a closed loop.
Select any service in the diagram to see its role. In words:
| Step | Service | Consumes | Produces |
|---|---|---|---|
| 1. Collect | k8s-monitor | Cluster APIs | MetricPoint every 30s |
| 2. Record | observability-agent | MetricPoint | TelemetrySnapshot → observability.telemetry |
| 3. Detect & decide | analysis-agent | observability.telemetry, feedback.signals | Decisions → analysis.decisions; insights → analysis.insights |
| 4. Act | action-agent | analysis.decisions | ActionOutcome → action.outcomes |
| 5. Measure | feedback-agent | action.outcomes, next snapshot | FeedbackSignal → feedback.signals |
| 6. Learn | analysis-agent | feedback.signals | Adjusted thresholds |
| — Advise | recommendation-agent | analysis.insights | Recommendations |
| — Investigate | agent-runtime | analysis.insights | An SRE investigation when risk > 70 |
A worked example
Suppose payments-api starts crash-looping after a bad config change.
- Collect. On the next cycle, k8s-monitor reports
CrashLoopCount: 3for the cluster. - Detect. The analysis agent compares 3 with its rolling baseline (normally 0). The Z-score is far above the critical threshold, so it records a critical anomaly and a restart decision — auto-approved, because restarting a crash-looping pod is safe.
- Act. The action agent deletes the failing pod; its Deployment recreates it. It publishes an
ActionOutcomewith statussuccess. - Measure. The feedback agent waits for the next snapshot.
CrashLoopCountis still 3 — the restart didn't help, because the config itself is wrong. It publishes a correction. - Learn. The analysis agent raises that cluster's crash-loop threshold slightly, so it won't keep restarting pods for this kind of failure.
- Escalate. Meanwhile the insight's risk score passed 70, so agent-runtime launched an SRE Orchestrator run. It reads the pod logs, sees the config error, links it to the deploy two minutes earlier from
cicd.events, and recommends rolling back.
You see all of this on the dashboard: the anomaly in Analytics, the action on /agents/actions, the feedback on /agents/feedback, and the investigation in Agent Runs.
Adaptive thresholds
Every cluster behaves differently: a spiky batch cluster and a steady web cluster shouldn't share one definition of "abnormal." So the analysis agent keeps a threshold per cluster, per metric, and tunes it from feedback:
| Signal | Meaning | Threshold change |
|---|---|---|
reinforcement | The action improved the affected metric. | −0.05 — detect sooner next time |
correction | The action failed, or didn't improve the metric. | +0.10 — be more cautious next time |
Thresholds stay between 1.5 and 5.0. Corrections move thresholds twice as fast as reinforcements, so the system becomes quieter quickly when its actions aren't helping, and more sensitive gradually when they are.
What you control
| Setting | Where | Effect |
|---|---|---|
| Which decisions run automatically | analysis-agent decision rules | Decide what runs without approval. |
| Approve or reject pending actions | /agents/actions, or the action API | Keep a human in the loop for risky actions. |
| Starting thresholds and window size | analysis-agent configuration | Tune initial sensitivity. |
| Scale-up verification | action-agent configuration | How long to wait and how much health may drop before reverting. |
| Auto-investigation threshold | agent-runtime configuration | The risk score that launches an SRE investigation. |
Message schemas
Telemetry
type TelemetryPublishMessage struct {
SnapshotID string
ClusterID string
Timestamp time.Time
Snapshot TelemetrySnapshot
}
type TelemetrySnapshot struct {
ClusterID string
Source string // "k8s-monitor" | "security-api" | "cicd-gateway"
HealthScore float64
CPUUsagePct float64
MemoryUsagePct float64
ReadyNodeCount int
NodeCount int
PodCount int
FailedPodCount int
CrashLoopCount int
APIServerLatencyMs float64
NetworkRxBytesPS float64
NetworkTxBytesPS float64
SecurityPostureScore float64
}
ClusterID identifies what was measured: a host cluster, or a tenant's vCluster. See observability-agent.
Decisions
| Type | Action taken by action-agent |
|---|---|
scale | Scale the target Deployment's replicas. |
restart | Delete the target Pod so its controller recreates it. |
cordon | Mark the target Node unschedulable. |
notify | Record and notify only; no Kubernetes change. |
Related: anomaly-detector alert rules
The anomaly detector is a second, rule-based remediation path. It consumes anomaly events from k8s.anomalies (published by k8s-monitor and the Optimizer) and evaluates alert rules you define — useful when you want an explicit, named rule for a specific workload:
POST /api/anomalies/api/v1/alert-rules
Content-Type: application/json
{
"cluster_id": "prod-us-east",
"name": "High CPU on API pods",
"metric": "cpu_usage_pct",
"condition": "z_score_gt",
"threshold": 2.5,
"severity": "high",
"action": "scale_up",
"enabled": true,
"auto_approve": false
}
With auto_approve: true the action runs automatically; with false the anomaly is recorded and shown for a person to act on.
Next steps
- analysis-agent — detection, risk scoring and decision rules.
- action-agent — every action, its safety checks and approvals.
- feedback-agent — how outcomes are measured.