Skip to main content
Version: 1.0

Agentic AI Layer

KubeOpera runs two complementary AI systems: a reactive event-driven pipeline and a Claude-powered reasoning runtime. Together they handle everything from automated remediation to deep investigative analysis.

Agent Mesh​

Agent mesh

The reactive pipeline's agents and how they work together.

observability.telemetry
decisions
feedback.signals
insights
Select any agent to see its role. Action outcomes flow to the feedback agent, which tunes the analysis agent.

Two AI Systems​

Reactive PipelineReasoning Runtime
TriggerEvery telemetry cycle (30s)Manual or auto (RiskScore > 70)
Latency< 2s end-to-end30s–5 min per run
ModelRule-based (Z-score) + Claude for recommendationsClaude with extended thinking + tool use
ActionsKubernetes remediations (restart, cordon, scale)Investigation + prioritised action list
OutputActionOutcome + FeedbackSignal + RecommendationsAgentRun with full tool call transcript
MemoryAdaptive thresholds per cluster per metric — built, but not wired up yet (see below)Per-run context; persistent run history

Reactive Pipeline​

The reactive pipeline is a chain of five specialised microservices connected by RabbitMQ topic exchanges:

  1. k8s-monitor collects MetricPoint every 30 seconds from all registered clusters
  2. observability-agent-srv persists TelemetrySnapshot and publishes to observability.telemetry
  3. analysis-agent-srv runs Z-score detection, produces AnalysisResult with RiskScore
  4. action-agent-srv executes Kubernetes actions for AutoApprove: true decisions
  5. feedback-agent-srv compares post-action metrics to the pre-action baseline and publishes a FeedbackSignal (reinforcement or correction)

recommendation-agent-srv and agent-runtime consume analysis.insights in parallel — one generating Claude-powered recommendations, the other auto-launching SRE investigations when risk is high.

The loop this diagram implies — feedback tuning analysis's own thresholds — isn't actually closed today. analysis-agent-srv has real, tested logic to adjust its Z-score thresholds (+0.1 on a correction, −0.05 on a reinforcement) from feedback.signals, but it never starts a consumer for that exchange, so the adjustment code has never run in practice. Both halves are real; only the wiring between them is missing. See Analysis Agent for detail.

Reasoning Runtime​

The agent-runtime service runs Claude-powered agents using the ReAct pattern (Reason → Act → Observe → repeat). Its full tool set spans Kubernetes inspection, observability queries, incident/runbook actions, cost data, and security posture — see Agentic Runtime for the exact tool count and per-agent breakdown, which has changed enough times across this service's life that it's worth trusting only that one page rather than duplicating a number here.

Seven agent types are available today: sre_orchestrator, security_auditor, cost_optimizer, incident_responder, node_ops, load_test_analyst, and app_advisor. Each is scoped to a subset of the full tool set for its domain, except SRE Orchestrator, which gets all of them. The dashboard's own "New Run" modal currently only offers four of these seven from its dropdown (sre_orchestrator, security_auditor, cost_optimizer, incident_responder) — the other three can be triggered via a direct API call.

RabbitMQ Exchanges​

ExchangeProducerConsumers
observability.telemetryobservability-agent-srvanalysis-agent-srv
analysis.decisionsanalysis-agent-srvaction-agent-srv
analysis.insightsanalysis-agent-srvrecommendation-agent-srv, agent-runtime
action.outcomesaction-agent-srvfeedback-agent-srv
feedback.signalsfeedback-agent-srvnone today — analysis-agent-srv never starts a consumer for this exchange, so messages here are published but never read
security.posturesecurity-apiincident-manager
cicd.eventscicd-gatewaynone today — a real publisher exists, but nothing in the platform subscribes to it yet