Skip to main content
Version: 1.0

Components

The platform consists of the following core components and services:

ComponentDescription
Control PlaneAPI Gateway, orchestration engine, and management interface
Build PipelineCI/CD system for building and testing container images
Registry ServiceContainer image storage and distribution
Cluster ManagerKubernetes cluster provisioning and lifecycle management
Deployment EngineApplication deployment orchestration and rollout management
Observability StackMonitoring, logging, and metrics collection
Security LayerAuthentication, authorization, secrets management, and policy enforcement

Management Interface​

KubeOpera-UI​

The primary user interface. Built with Next.js 14 App Router, TypeScript, and Tailwind CSS.

Key pages and features:

  • Dashboard — cluster health overview with live metric cards
  • Clusters — provision and manage clusters; multi-cluster comparison treemap
  • Pipelines — CI/CD run history with stage timeline charts (recharts)
  • Analytics — forecast charts with confidence bands, anomaly table, scaling approvals
  • Incidents — two-panel incident list + detail with vertical timeline
  • AI Agents — agent health overview (live polling), event flow diagram
  • Agent Runs — list of reasoning agent runs; live SSE streaming viewer with thinking blocks and tool call trace
  • Monitoring — cost, performance, and security monitoring pages
  • Node Pools — Karpenter NodePool and NodeClaim viewer with spot market pricing table
  • Monitoring — SLOs (error budget + burn rate gauges), APM (service RED metrics + slowest/erroring endpoints), Traces (Jaeger service map + slow/error trace explorer with span detail)
  • AI Chat — streaming chat with 30 tool integrations across all services

All backend communication goes through Next.js API proxy routes (/api/*) that enforce the SSRF allowlist and forward requests to the appropriate microservice.

Infrastructure Services​

K8s-Monitor​

Polls the Kubernetes API and cloud provider APIs on a configurable cycle. Produces:

  • Cluster health score (0–100) based on node readiness, pod health, and control plane status
  • Per-namespace and per-workload cost breakdowns (AWS, GCP, on-prem estimation)
  • Resource optimisation recommendations (over-provisioned, idle workloads)
  • Pod and node metrics

Kubeopera-API​

Core application and cluster management API. Handles:

  • Cluster registration and lifecycle
  • Application (workload) CRUD
  • App creation triggers KubeOperaApp CRD creation (see app-controller)

Security Service​

Three-service scanning pipeline:

  • security-cron — scheduled trigger; initiates scan jobs
  • security-collector — executes vulnerability and configuration scans against cluster resources; publishes findings to RabbitMQ security.posture exchange
  • security-api — stores findings in PostgreSQL; serves posture summary, findings, and compliance score

Nodes Manager​

Intelligent node lifecycle management. Handles Karpenter NodePool/NodeClaim control, AI-powered workload scheduling (5-dimension scoring), AWS spot market analysis, and proactive consolidation planning. Syncs node state from Kubernetes to PostgreSQL on a background cycle and publishes scaling events to the nodes.events RabbitMQ exchange.

Auth Service​

Handles user authentication, token issuance, and session management for the platform. Integrates with OIDC providers and maintains local user accounts.

Cache Service​

Redis-backed caching layer used by kubeopera-api and other services to reduce database load on frequently read data.

Automation Services​

CICD Service​

Receives GitHub and GitLab webhooks (HMAC-SHA256 verified), persists pipeline runs, and publishes events to the cicd.events RabbitMQ exchange.

Domain models: Pipeline, PipelineRun (with status: queued/running/success/failed/cancelled), StageRun

Anomaly Detector​

Consumes anomaly events from k8s.anomalies exchange (published by k8s-monitor when Z-score thresholds are crossed). For each event:

  1. Evaluates AlertRule definitions (per-cluster, per-metric)
  2. Executes self-healing actions if AutoApprove: true (pod restart / node cordon / scale deployment)
  3. Logs the RemediationLog entry
  4. Publishes to k8s.selfheal exchange

Predictive Scaler​

Reads historical metric snapshots from PostgreSQL. Runs double exponential smoothing (Holt-Winters) in pure Go to generate load forecasts with ±1.96×RMSE confidence bands. Produces ScalingDecision records that can be approved from the Analytics UI or via the AI agent.

Requires a minimum of 24 data points before generating forecasts. No external ML libraries.

Incident Manager​

Full incident lifecycle service:

  • Consumes k8s.anomalies, k8s.selfheal, and security.posture exchanges
  • Correlates events using a 5-minute sliding window keyed on {clusterID}:{category}:{namespace} to prevent alert storms
  • Executes runbooks (kubectl / HTTP / notify / wait steps)
  • Sends notifications to Slack webhooks and PagerDuty
  • Tracks MTTR, MTBF, and SLA breach percentages

App Lifecycle Services​

App Service​

Manifest generation engine. Takes a schema definition (app name, image, ports, resources, autoscaling, ingress) and generates production-ready Kubernetes YAML: Deployment, Service, Ingress, and the rest of what an app needs to actually run. Exposed as an HTTP preview endpoint (POST /api/v1/preview) that App Controller calls internally.

App Controller​

The Kubernetes operator that reconciles KubeOperaApp custom resources — see its own service page for the full reconciliation flow, including how it resolves the correct tenant vCluster and handles finalizer-based cascade deletion. In outline: it generates manifests via App Service, pushes them to the fleet Git repository, and ensures Flux GitRepository/Kustomization resources exist inside the target vCluster to pick the change up.

Agentic AI Layer​

Observability Agent​

Collects telemetry from k8s-monitor on a configurable interval. Publishes TelemetrySnapshot messages to the observability.telemetry RabbitMQ exchange. Stores snapshots in PostgreSQL for historical queries.

Analysis Agent​

The statistical brain of the pipeline. Consumes telemetry from observability.telemetry and:

  • Runs Z-score anomaly detection over a 60-point rolling window (default threshold: 2.5 high / 3.5 critical)
  • Produces AnalysisResult with decisions (scale / restart / cordon / notify) and risk score
  • Publishes DecisionPublishMessage to analysis.decisions and InsightPublishMessage to analysis.insights
  • Has real, working logic to adapt its Z-score thresholds (+0.1 on correction, -0.05 on reinforcement) from feedback.signals — but as of this writing, it never actually starts a consumer for that exchange, so this adaptive loop has never run in practice despite both halves of it (Feedback Agent's publish, Analysis Agent's threshold-adjustment logic) being real, tested code

Action Agent​

Consumes analysis.decisions and executes Kubernetes remediation actions:

  • scale_deployment — increments replica count by 1 (up to HPA maxReplicas)
  • restart_pod — deletes pod, allowing the controller to recreate it
  • cordon_node — patches spec.unschedulable: true

Only processes decisions with AutoApprove: true. All actions produce ActionOutcome events published to action.outcomes.

Feedback Agent​

Evaluates whether each remediation action executed successfully (not, as the name might suggest, whether it actually improved cluster state via a follow-up telemetry comparison — there is no such comparison today). Publishes a FeedbackSignal (type: reinforcement on success, correction on failure) to the feedback.signals exchange — intended to let the analysis agent tune its Z-score thresholds, though nothing there currently consumes it.

Recommendation Agent​

Consumes analysis.insights and calls the Anthropic API (Claude Sonnet 4.5) to generate structured recommendations (category, priority, title, description, impact, action). Stores recommendations in PostgreSQL and serves them via REST.

Agentic Runtime​

The ReAct reasoning engine. Uses BetaToolRunnerStreaming from anthropic-sdk-go v1.37 to run seven agent types across 34 tools:

AgentModelToolsMax iterations
SRE OrchestratorClaude Sonnet 4.6All 3420
Security AuditorClaude Haiku 4.55 (security)10
Cost OptimizerClaude Haiku 4.55 (cost)10
Incident ResponderClaude Sonnet 4.613 (incident + write)15
App AdvisorClaude Sonnet 4.67 (app advisor)12
Load Test AnalystClaude Sonnet 4.67 (performance)10
NodeOpsClaude Sonnet 4.610 (node management)15

SRE Orchestrator, Incident Responder, and NodeOps use interleaved thinking (interleaved-thinking-2025-05-14 beta). Every tool call is wrapped by instrumentedTool which emits SSE events and persists to the agents.tool_calls table. Auto-triggered when analysis risk score exceeds 70.

Observability Services​

Log Service​

Stateless Loki façade. Translates service/namespace queries into LogQL and returns log lines and extracted patterns. No database.

SLO Manager​

Stores SLO definitions and evaluates them against live Prometheus queries. Computes error budget (remaining fraction) and burn rate for each SLO.

APM Gateway​

Stateless Prometheus façade for application performance monitoring. Provides per-service request rates (RPS), p99 latency, and 5xx error rate via standard PromQL queries.

Tracing Gateway​

Stateless Jaeger HTTP API v3 façade. Returns trace detail, service dependency map, and search results for slow and error traces.

RCA Engine​

AI-powered root cause analysis. Parallelises signal collection from anomaly-detector, log-gateway, and incident-manager, then calls Claude Haiku to produce a structured RCA with root cause, contributing factors, and confidence score.

Optimiser​

Continuous AI optimisation agent powered by Claude. Runs a background monitoring loop that assesses cluster health and exposes HTTP endpoints for on-demand health analysis, resource optimisation, load test analysis, and interactive kubectl/Prometheus troubleshooting. The get_optimization_status and get_load_test_analysis tools connect it to agent-runtime, the AI chat, and the MCP server. It also publishes degradation signals to the k8s.anomalies exchange, but a routing-key mismatch means anomaly-detector never actually receives them today — see Optimizer for the full detail.

MCP Server​

A fork of containers/kubernetes-mcp-server with a custom kubeopera toolset plugin — 49 tools in total, across read-only data tools, two write action tools (drain_node_action, rollback_deployment), six agent launchers, and dedicated App Advisor, observability, and GitOps toolsets. See MCP Server for the full breakdown. Integrates with Claude Desktop, Claude Code, and any MCP-compatible client.