Skip to main content
Version: 2.0

Core Concepts

This page introduces the ideas KubeOpera is built on and the vocabulary used throughout the documentation. Once these click, the rest of the platform will feel predictable.

Why KubeOpera exists​

Running Kubernetes well usually means assembling and maintaining a toolchain: a provisioning tool, a GitOps controller, a metrics stack, a log stack, a tracing stack, a cost tool, a security scanner, an incident tool, and a lot of glue. Each works; together they are a second system to operate.

KubeOpera is built on Kubernetes' own primitives and on proven open-source tools — Flux, Prometheus, Loki, Jaeger, Karpenter, vCluster — and wires them into one platform with one data model, one API and one UI. On top of that foundation it adds an AI layer that can reason about everything the platform sees.

Design principles​

Event-driven by default​

Services talk to each other through RabbitMQ topic exchanges. When telemetry is collected, the services that care about it receive it immediately and act — there is no polling between services and no central coordinator. New capabilities are added by subscribing to existing events, without changing the services that publish them.

Autonomous, with humans in control​

KubeOpera can restart pods, cordon nodes and scale workloads on its own. Every automated action is governed by an approval rule (AutoApprove), logged with its reason, and measured afterwards. You decide which actions run automatically and which wait for a person to approve them.

A closed loop, not a one-way pipe​

Detection → decision → action → feedback. After every automated action KubeOpera checks whether cluster health actually improved, and feeds that result back into detection. Over time each cluster gets thresholds tuned to its own behaviour: fewer false alarms where actions don't help, faster alerts where they do.

AI as a first-class operator​

Reasoning agents built on Claude use live, typed tools to inspect clusters, correlate signals and propose — or, when permitted, take — action. They can be launched by a person, by another service, or automatically when risk rises.

Transparent AI​

Every step an agent takes — its reasoning, each tool call and each result — is streamed to the UI as it happens and stored for later review. Nothing an agent does is hidden.

Declarative and GitOps-native​

Applications, tenants and clusters are described as Kubernetes custom resources and reconciled by controllers. Application manifests land in Git and are delivered by Flux, so every change is reviewable, auditable and revertible.

Building blocks​

ConceptWhat it is
Host clusterA Kubernetes cluster that KubeOpera manages. It runs tenant virtual clusters and application workloads.
Management planeThe KubeOpera services themselves — APIs, controllers, agents — running in the kubeopera-* namespaces.
TenantAn organization or team using KubeOpera. Each tenant is isolated from the others.
CloudSpace / vClusterA tenant's own virtual Kubernetes cluster, running inside a host cluster, with its own API server and quotas.
KubeOperaAppThe custom resource that describes an application. The app-controller turns it into manifests in Git that Flux deploys.
Telemetry snapshotA point-in-time record of a cluster's health, resource usage and failures, collected every 30 seconds.
DecisionA proposed remediation (scale, restart, cordon, notify) produced by analysis, with an approval rule.
Feedback signalThe measured outcome of an action — reinforcement if it helped, correction if it didn't — used to tune detection.
Agent runOne session of a reasoning agent working on a goal, with its full reasoning and tool-call history.
RecommendationA structured, prioritized suggestion (cost, reliability, security, performance) generated by AI.

Key capabilities​

AreaWhat it does
Multi-cluster managementProvision, register and operate clusters across AWS, GCP and on-premises from one UI.
Multi-tenancyIsolated virtual clusters per tenant with quotas and RBAC.
Application deliverySchema-driven manifest generation and GitOps deployment, from source or image.
ObservabilityHealth scores, metrics, logs, traces, SLOs and APM, correlated per workload.
Anomaly detectionStatistical detection over rolling windows with adaptive, per-cluster thresholds.
Self-healingAutomated restarts, cordons and scale-outs, each measured for effect.
Predictive scalingLoad forecasts with confidence bands and one-click scaling approvals.
Cost intelligenceReal-time cost per namespace and workload, forecasts and right-sizing.
Security postureContinuous scanning, RBAC analysis and compliance scoring.
Incident managementCorrelated incidents, runbooks, notifications and reliability metrics.
Agentic AIReasoning agents with live tools, streaming and full audit trails.

Technology stack​

LayerTechnologyWhy
DashboardNext.js, TypeScript, Tailwind CSSServer components, streaming UI, a server-side proxy for every backend call.
ServicesGo, hexagonal architectureFast, small, single-binary services with clear boundaries.
MessagingRabbitMQ topic exchangesFan-out without coupling producers to consumers.
StoragePostgreSQL, RedisDurable state and fast caching.
AIClaude via the Anthropic Go SDKTool use, streaming and extended thinking.
ForecastingHolt-Winters, implemented in GoAccurate short-term forecasts with no ML framework dependencies.
DeliveryFlux v2 via controller-runtimeDeclarative, Kubernetes-native reconciliation.
TenancyvClusterStrong isolation without a cluster per team.
NodesKarpenterFast, cost-aware node provisioning, including spot capacity.

Next steps​

  • Architecture — see how the layers and services fit together.
  • Components — a tour of every major component and what it is responsible for.
  • AI Agents — how the autonomous loop and reasoning agents work.