Agentic AI Layer
KubeOpera runs two complementary AI systems: a reactive event-driven pipeline and a Claude-powered reasoning runtime. Together they handle everything from automated remediation to deep investigative analysis.
Agent Mesh
Agent mesh
The reactive pipeline's agents and how they work together.
Two AI Systems
| Reactive Pipeline | Reasoning Runtime | |
|---|---|---|
| Trigger | Every telemetry cycle (30s) | Manual or auto (RiskScore > 70) |
| Latency | < 2s end-to-end | 30s–5 min per run |
| Model | Rule-based (Z-score) + Claude for recommendations | Claude with extended thinking + tool use |
| Actions | Kubernetes remediations (restart, cordon, scale) | Investigation + prioritised action list |
| Output | ActionOutcome + FeedbackSignal + Recommendations | AgentRun with full tool call transcript |
| Memory | Adaptive thresholds per cluster per metric — built, but not wired up yet (see below) | Per-run context; persistent run history |
Reactive Pipeline
The reactive pipeline is a chain of five specialised microservices connected by RabbitMQ topic exchanges:
- k8s-monitor collects MetricPoint every 30 seconds from all registered clusters
- observability-agent-srv persists TelemetrySnapshot and publishes to
observability.telemetry - analysis-agent-srv runs Z-score detection, produces AnalysisResult with RiskScore
- action-agent-srv executes Kubernetes actions for
AutoApprove: truedecisions - feedback-agent-srv compares post-action metrics to the pre-action baseline and publishes a
FeedbackSignal(reinforcement or correction)
recommendation-agent-srv and agent-runtime consume analysis.insights in parallel — one generating Claude-powered recommendations, the other auto-launching SRE investigations when risk is high.
The loop this diagram implies — feedback tuning analysis's own thresholds — isn't actually closed today. analysis-agent-srv has real, tested logic to adjust its Z-score thresholds (+0.1 on a correction, −0.05 on a reinforcement) from feedback.signals, but it never starts a consumer for that exchange, so the adjustment code has never run in practice. Both halves are real; only the wiring between them is missing. See Analysis Agent for detail.
Reasoning Runtime
The agent-runtime service runs Claude-powered agents using the ReAct pattern (Reason → Act → Observe → repeat). Its full tool set spans Kubernetes inspection, observability queries, incident/runbook actions, cost data, and security posture — see Agentic Runtime for the exact tool count and per-agent breakdown, which has changed enough times across this service's life that it's worth trusting only that one page rather than duplicating a number here.
Seven agent types are available today: sre_orchestrator, security_auditor, cost_optimizer, incident_responder, node_ops, load_test_analyst, and app_advisor. Each is scoped to a subset of the full tool set for its domain, except SRE Orchestrator, which gets all of them. The dashboard's own "New Run" modal currently only offers four of these seven from its dropdown (sre_orchestrator, security_auditor, cost_optimizer, incident_responder) — the other three can be triggered via a direct API call.
RabbitMQ Exchanges
| Exchange | Producer | Consumers |
|---|---|---|
observability.telemetry | observability-agent-srv | analysis-agent-srv |
analysis.decisions | analysis-agent-srv | action-agent-srv |
analysis.insights | analysis-agent-srv | recommendation-agent-srv, agent-runtime |
action.outcomes | action-agent-srv | feedback-agent-srv |
feedback.signals | feedback-agent-srv | none today — analysis-agent-srv never starts a consumer for this exchange, so messages here are published but never read |
security.posture | security-api | incident-manager |
cicd.events | cicd-gateway | none today — a real publisher exists, but nothing in the platform subscribes to it yet |