Components
The platform consists of the following core components and services:
| Component | Description |
|---|---|
| Control Plane | API Gateway, orchestration engine, and management interface |
| Build Pipeline | CI/CD system for building and testing container images |
| Registry Service | Container image storage and distribution |
| Cluster Manager | Kubernetes cluster provisioning and lifecycle management |
| Deployment Engine | Application deployment orchestration and rollout management |
| Observability Stack | Monitoring, logging, and metrics collection |
| Security Layer | Authentication, authorization, secrets management, and policy enforcement |
Management Interface
KubeOpera-UI
The primary user interface. Built with Next.js 14 App Router, TypeScript, and Tailwind CSS.
Key pages and features:
- Dashboard — cluster health overview with live metric cards
- Clusters — provision and manage clusters; multi-cluster comparison treemap
- Pipelines — CI/CD run history with stage timeline charts (recharts)
- Analytics — forecast charts with confidence bands, anomaly table, scaling approvals
- Incidents — two-panel incident list + detail with vertical timeline
- AI Agents — agent health overview (live polling), event flow diagram
- Agent Runs — list of reasoning agent runs; live SSE streaming viewer with thinking blocks and tool call trace
- Monitoring — cost, performance, and security monitoring pages
- Node Pools — Karpenter NodePool and NodeClaim viewer with spot market pricing table
- Monitoring — SLOs (error budget + burn rate gauges), APM (service RED metrics + slowest/erroring endpoints), Traces (Jaeger service map + slow/error trace explorer with span detail)
- AI Chat — streaming chat with 30 tool integrations across all services
All backend communication goes through Next.js API proxy routes (/api/*) that enforce the SSRF allowlist and forward requests to the appropriate microservice.
Infrastructure Services
K8s-Monitor
Polls the Kubernetes API and cloud provider APIs on a configurable cycle. Produces:
- Cluster health score (0–100) based on node readiness, pod health, and control plane status
- Per-namespace and per-workload cost breakdowns (AWS, GCP, on-prem estimation)
- Resource optimisation recommendations (over-provisioned, idle workloads)
- Pod and node metrics
Kubeopera-API
Core application and cluster management API. Handles:
- Cluster registration and lifecycle
- Application (workload) CRUD
- App creation triggers
KubeOperaAppCRD creation (see app-controller)
Security Service
Three-service scanning pipeline:
security-cron— scheduled trigger; initiates scan jobssecurity-collector— executes vulnerability and configuration scans against cluster resources; publishes findings to RabbitMQsecurity.postureexchangesecurity-api— stores findings in PostgreSQL; serves posture summary, findings, and compliance score
Nodes Manager
Intelligent node lifecycle management. Handles Karpenter NodePool/NodeClaim control, AI-powered workload scheduling (5-dimension scoring), AWS spot market analysis, and proactive consolidation planning. Syncs node state from Kubernetes to PostgreSQL on a background cycle and publishes scaling events to the nodes.events RabbitMQ exchange.
Auth Service
Handles user authentication, token issuance, and session management for the platform. Integrates with OIDC providers and maintains local user accounts.
Cache Service
Redis-backed caching layer used by kubeopera-api and other services to reduce database load on frequently read data.
Automation Services
CICD Service
Receives GitHub and GitLab webhooks (HMAC-SHA256 verified), persists pipeline runs, and publishes events to the cicd.events RabbitMQ exchange.
Domain models: Pipeline, PipelineRun (with status: queued/running/success/failed/cancelled), StageRun
Anomaly Detector
Consumes anomaly events from k8s.anomalies exchange (published by k8s-monitor when Z-score thresholds are crossed). For each event:
- Evaluates
AlertRuledefinitions (per-cluster, per-metric) - Executes self-healing actions if
AutoApprove: true(pod restart / node cordon / scale deployment) - Logs the
RemediationLogentry - Publishes to
k8s.selfhealexchange
Predictive Scaler
Reads historical metric snapshots from PostgreSQL. Runs double exponential smoothing (Holt-Winters) in pure Go to generate load forecasts with ±1.96×RMSE confidence bands. Produces ScalingDecision records that can be approved from the Analytics UI or via the AI agent.
Requires a minimum of 24 data points before generating forecasts. No external ML libraries.
Incident Manager
Full incident lifecycle service:
- Consumes
k8s.anomalies,k8s.selfheal, andsecurity.postureexchanges - Correlates events using a 5-minute sliding window keyed on
{clusterID}:{category}:{namespace}to prevent alert storms - Executes runbooks (kubectl / HTTP / notify / wait steps)
- Sends notifications to Slack webhooks and PagerDuty
- Tracks MTTR, MTBF, and SLA breach percentages
App Lifecycle Services
App Service
Manifest generation engine. Takes a schema definition (app name, image, ports, resources, autoscaling, ingress) and generates production-ready Kubernetes YAML: Deployment, Service, Ingress, and the rest of what an app needs to actually run. Exposed as an HTTP preview endpoint (POST /api/v1/preview) that App Controller calls internally.
App Controller
The Kubernetes operator that reconciles KubeOperaApp custom resources — see its own service page for the full reconciliation flow, including how it resolves the correct tenant vCluster and handles finalizer-based cascade deletion. In outline: it generates manifests via App Service, pushes them to the fleet Git repository, and ensures Flux GitRepository/Kustomization resources exist inside the target vCluster to pick the change up.
Agentic AI Layer
Observability Agent
Collects telemetry from k8s-monitor on a configurable interval. Publishes TelemetrySnapshot messages to the observability.telemetry RabbitMQ exchange. Stores snapshots in PostgreSQL for historical queries.
Analysis Agent
The statistical brain of the pipeline. Consumes telemetry from observability.telemetry and:
- Runs Z-score anomaly detection over a 60-point rolling window (default threshold: 2.5 high / 3.5 critical)
- Produces
AnalysisResultwith decisions (scale / restart / cordon / notify) and risk score - Publishes
DecisionPublishMessagetoanalysis.decisionsandInsightPublishMessagetoanalysis.insights - Has real, working logic to adapt its Z-score thresholds (+0.1 on correction, -0.05 on reinforcement) from
feedback.signals— but as of this writing, it never actually starts a consumer for that exchange, so this adaptive loop has never run in practice despite both halves of it (Feedback Agent's publish, Analysis Agent's threshold-adjustment logic) being real, tested code
Action Agent
Consumes analysis.decisions and executes Kubernetes remediation actions:
scale_deployment— increments replica count by 1 (up to HPA maxReplicas)restart_pod— deletes pod, allowing the controller to recreate itcordon_node— patchesspec.unschedulable: true
Only processes decisions with AutoApprove: true. All actions produce ActionOutcome events published to action.outcomes.
Feedback Agent
Evaluates whether each remediation action executed successfully (not, as the name might suggest, whether it actually improved cluster state via a follow-up telemetry comparison — there is no such comparison today). Publishes a FeedbackSignal (type: reinforcement on success, correction on failure) to the feedback.signals exchange — intended to let the analysis agent tune its Z-score thresholds, though nothing there currently consumes it.
Recommendation Agent
Consumes analysis.insights and calls the Anthropic API (Claude Sonnet 4.5) to generate structured recommendations (category, priority, title, description, impact, action). Stores recommendations in PostgreSQL and serves them via REST.
Agentic Runtime
The ReAct reasoning engine. Uses BetaToolRunnerStreaming from anthropic-sdk-go v1.37 to run seven agent types across 34 tools:
| Agent | Model | Tools | Max iterations |
|---|---|---|---|
| SRE Orchestrator | Claude Sonnet 4.6 | All 34 | 20 |
| Security Auditor | Claude Haiku 4.5 | 5 (security) | 10 |
| Cost Optimizer | Claude Haiku 4.5 | 5 (cost) | 10 |
| Incident Responder | Claude Sonnet 4.6 | 13 (incident + write) | 15 |
| App Advisor | Claude Sonnet 4.6 | 7 (app advisor) | 12 |
| Load Test Analyst | Claude Sonnet 4.6 | 7 (performance) | 10 |
| NodeOps | Claude Sonnet 4.6 | 10 (node management) | 15 |
SRE Orchestrator, Incident Responder, and NodeOps use interleaved thinking (interleaved-thinking-2025-05-14 beta). Every tool call is wrapped by instrumentedTool which emits SSE events and persists to the agents.tool_calls table. Auto-triggered when analysis risk score exceeds 70.
Observability Services
Log Service
Stateless Loki façade. Translates service/namespace queries into LogQL and returns log lines and extracted patterns. No database.
SLO Manager
Stores SLO definitions and evaluates them against live Prometheus queries. Computes error budget (remaining fraction) and burn rate for each SLO.
APM Gateway
Stateless Prometheus façade for application performance monitoring. Provides per-service request rates (RPS), p99 latency, and 5xx error rate via standard PromQL queries.
Tracing Gateway
Stateless Jaeger HTTP API v3 façade. Returns trace detail, service dependency map, and search results for slow and error traces.
RCA Engine
AI-powered root cause analysis. Parallelises signal collection from anomaly-detector, log-gateway, and incident-manager, then calls Claude Haiku to produce a structured RCA with root cause, contributing factors, and confidence score.
Optimiser
Continuous AI optimisation agent powered by Claude. Runs a background monitoring loop that assesses cluster health and exposes HTTP endpoints for on-demand health analysis, resource optimisation, load test analysis, and interactive kubectl/Prometheus troubleshooting. The get_optimization_status and get_load_test_analysis tools connect it to agent-runtime, the AI chat, and the MCP server. It also publishes degradation signals to the k8s.anomalies exchange, but a routing-key mismatch means anomaly-detector never actually receives them today — see Optimizer for the full detail.
MCP Server
A fork of containers/kubernetes-mcp-server with a custom kubeopera toolset plugin — 49 tools in total, across read-only data tools, two write action tools (drain_node_action, rollback_deployment), six agent launchers, and dedicated App Advisor, observability, and GitOps toolsets. See MCP Server for the full breakdown. Integrates with Claude Desktop, Claude Code, and any MCP-compatible client.