AI Agents
KubeOpera runs two complementary AI systems:
- A reactive pipeline — small, fast services that watch every cluster continuously, fix known problems automatically, and learn from the results.
- A reasoning runtime — Claude-powered agents that investigate open-ended questions step by step, using live tools, and explain what they find.
The reactive pipeline is your always-on first responder; the reasoning runtime is the expert you call in when a problem needs thought. They share data and hand work to each other.
Agent mesh
Agent mesh
The reactive pipeline's agents and how they work together.
Two systems at a glance
| Reactive pipeline | Reasoning runtime | |
|---|---|---|
| Triggered by | Every telemetry cycle (30 seconds) | A person, another service, or automatically when risk exceeds 70 |
| Speed | Under 2 seconds end to end | 30 seconds to a few minutes per run |
| Intelligence | Statistical detection (adaptive Z-scores), plus Claude for recommendations | Claude with extended thinking and tool use |
| Acts by | Kubernetes remediations: scale, restart, cordon | Investigating and proposing — or, when permitted, taking — prioritized actions |
| Produces | Action outcomes, feedback signals, recommendations | An agent run with findings and a full tool-call history |
| Learns from | Measured outcomes of its own actions, per cluster and metric | Run history and context supplied for each run |
The reactive pipeline
Five specialized services, connected by RabbitMQ topic exchanges, form a closed loop:
- k8s-monitor collects a
MetricPointfrom every cluster every 30 seconds. - observability-agent stores a telemetry snapshot and publishes it to
observability.telemetry. - analysis-agent detects anomalies, scores risk and decides what to do.
- action-agent carries out decisions that are approved to run automatically.
- feedback-agent measures whether each action helped, and tells analysis — which tunes its thresholds.
Alongside the loop, recommendation-agent turns insights into recommendations and agent-runtime launches an investigation when risk is high.
Read Reactive AI Pipeline for a full walkthrough.
The reasoning runtime
agent-runtime runs agents using the ReAct pattern — reason, act, observe, repeat. Each step the agent thinks about what it knows, calls a tool to learn more (or to act), reads the result, and decides what to do next, until it can answer.
KubeOpera includes seven specialized agents:
| Agent | Best for |
|---|---|
sre_orchestrator | Open-ended investigation with access to every tool. |
security_auditor | Security posture, findings and RBAC. |
cost_optimizer | Spend, waste, right-sizing and scaling decisions. |
incident_responder | Triage and remediation of active incidents. |
node_ops | Node pools, capacity and consolidation. |
load_test_analyst | Performance and load-test results. |
app_advisor | A single application's configuration and health. |
All seven can be launched from the dashboard or the API. Each is scoped to the tools for its domain; the SRE Orchestrator can use all of them.
Human control
Automation in KubeOpera is always governed:
- Approval rules. Every decision carries
AutoApprove. Actions that aren't auto-approved wait in an approval queue for a person (or an authorized agent) to approve or reject. - Guardrails. Actions respect PodDisruptionBudgets, and scale-ups are verified afterwards and reverted automatically if health drops.
- Audit. Every action, decision, tool call and agent step is recorded with its reason and outcome.
RabbitMQ exchanges
| Exchange | Producer | Consumers |
|---|---|---|
observability.telemetry | observability-agent-srv | analysis-agent-srv |
analysis.decisions | analysis-agent-srv | action-agent-srv |
analysis.insights | analysis-agent-srv | recommendation-agent-srv, agent-runtime |
action.outcomes | action-agent-srv | feedback-agent-srv |
feedback.signals | feedback-agent-srv | analysis-agent-srv |
security.posture | security-collector | incident-manager |
cicd.events | cicd-gateway | observability-agent-srv, incident-manager |
Deployment events from cicd.events let the pipeline tell "something broke" from "something broke right after a deploy" — and let incidents link straight to the change that likely caused them.
Next steps
- Reactive AI Pipeline — the loop in depth.
- Agent runtime — agents, tools and how runs work.
- Your first AI investigation — try it yourself.