Skip to main content
Version: 2.0

AI Agents

KubeOpera runs two complementary AI systems:

  • A reactive pipeline — small, fast services that watch every cluster continuously, fix known problems automatically, and learn from the results.
  • A reasoning runtime — Claude-powered agents that investigate open-ended questions step by step, using live tools, and explain what they find.

The reactive pipeline is your always-on first responder; the reasoning runtime is the expert you call in when a problem needs thought. They share data and hand work to each other.

Agent mesh​

Agent mesh

The reactive pipeline's agents and how they work together.

observability.telemetry
decisions
feedback.signals
insights
Select any agent to see its role. Action outcomes flow to the feedback agent, which tunes the analysis agent.

Two systems at a glance​

Reactive pipelineReasoning runtime
Triggered byEvery telemetry cycle (30 seconds)A person, another service, or automatically when risk exceeds 70
SpeedUnder 2 seconds end to end30 seconds to a few minutes per run
IntelligenceStatistical detection (adaptive Z-scores), plus Claude for recommendationsClaude with extended thinking and tool use
Acts byKubernetes remediations: scale, restart, cordonInvestigating and proposing — or, when permitted, taking — prioritized actions
ProducesAction outcomes, feedback signals, recommendationsAn agent run with findings and a full tool-call history
Learns fromMeasured outcomes of its own actions, per cluster and metricRun history and context supplied for each run

The reactive pipeline​

Five specialized services, connected by RabbitMQ topic exchanges, form a closed loop:

  1. k8s-monitor collects a MetricPoint from every cluster every 30 seconds.
  2. observability-agent stores a telemetry snapshot and publishes it to observability.telemetry.
  3. analysis-agent detects anomalies, scores risk and decides what to do.
  4. action-agent carries out decisions that are approved to run automatically.
  5. feedback-agent measures whether each action helped, and tells analysis — which tunes its thresholds.

Alongside the loop, recommendation-agent turns insights into recommendations and agent-runtime launches an investigation when risk is high.

Read Reactive AI Pipeline for a full walkthrough.

The reasoning runtime​

agent-runtime runs agents using the ReAct pattern — reason, act, observe, repeat. Each step the agent thinks about what it knows, calls a tool to learn more (or to act), reads the result, and decides what to do next, until it can answer.

KubeOpera includes seven specialized agents:

AgentBest for
sre_orchestratorOpen-ended investigation with access to every tool.
security_auditorSecurity posture, findings and RBAC.
cost_optimizerSpend, waste, right-sizing and scaling decisions.
incident_responderTriage and remediation of active incidents.
node_opsNode pools, capacity and consolidation.
load_test_analystPerformance and load-test results.
app_advisorA single application's configuration and health.

All seven can be launched from the dashboard or the API. Each is scoped to the tools for its domain; the SRE Orchestrator can use all of them.

Human control​

Automation in KubeOpera is always governed:

  • Approval rules. Every decision carries AutoApprove. Actions that aren't auto-approved wait in an approval queue for a person (or an authorized agent) to approve or reject.
  • Guardrails. Actions respect PodDisruptionBudgets, and scale-ups are verified afterwards and reverted automatically if health drops.
  • Audit. Every action, decision, tool call and agent step is recorded with its reason and outcome.

RabbitMQ exchanges​

ExchangeProducerConsumers
observability.telemetryobservability-agent-srvanalysis-agent-srv
analysis.decisionsanalysis-agent-srvaction-agent-srv
analysis.insightsanalysis-agent-srvrecommendation-agent-srv, agent-runtime
action.outcomesaction-agent-srvfeedback-agent-srv
feedback.signalsfeedback-agent-srvanalysis-agent-srv
security.posturesecurity-collectorincident-manager
cicd.eventscicd-gatewayobservability-agent-srv, incident-manager

Deployment events from cicd.events let the pipeline tell "something broke" from "something broke right after a deploy" — and let incidents link straight to the change that likely caused them.

Next steps​