Skip to main content
Version: 2.0

Architecture

KubeOpera is made of two codebases that work together: a Next.js dashboard and a set of Go microservices running in Kubernetes. They communicate in three ways:

  • REST — for requests a person or another service makes and waits for.
  • RabbitMQ topic exchanges — for events that many services react to independently.
  • Kubernetes custom resources — for desired state that controllers reconcile.

This page walks through the system layer by layer, then follows a request and an event through it.

System diagram​

Select any component to see what it does.

High-Level System Architecture

Select any component to see what it does
User InterfaceWho works in KubeOpera
API GatewayNext.js server-side proxy
Central Management PlaneGo microservices and AI analytics
Application Services
Analytics Engine
Kubernetes-as-a-ServiceManaged clusters and tenant spaces
Control Plane
etcdcontroller-managercloud-controller-managerschedulerkube-apiservercrm
Work Plane
KubeSpace 1
BillCtlrResourceCtlr
AI Agents
vCluster
apikubeletetcd
NodePool 1 · 4 nodes
KubeSpace 2
BillCtlrResourceCtlr
AI Agents
vCluster
apikubeletetcd
NodePool 2 · 4 nodes
KubeSpace 3
BillCtlrResourceCtlr
AI Agents
vCluster
apikubeletetcd
NodePool 3 · 4 nodes
KubeSpace 4
BillCtlrResourceCtlr
AI Agents
vCluster
apikubeletetcd
NodePool 4 · 4 nodes
Platform ToolsDelivery, observability and identity
Select any component above to learn what it does.

The layers​

1. User interface​

Everyone — platform engineers, developers, SREs and FinOps — works in the same dashboard. Pages are server-rendered where possible and stream live data (agent runs, logs, health) where it matters.

2. API gateway​

The dashboard never calls a backend service directly from the browser. Every call goes to a Next.js route under /api/*, which:

  1. Authenticates the request using the session's JWT issued by auth-service.
  2. Rate-limits per endpoint to protect downstream services.
  3. Enforces an allowlist — only known service hostnames can be proxied to, which prevents server-side request forgery.
  4. Routes the request to the right service based on its path prefix.
  5. Load-balances across service replicas through Kubernetes Services.

Keeping backend URLs and credentials on the server also means no CORS configuration and no secrets in the browser.

3. Central management plane​

This is the heart of the platform: Go services, each owning one domain (clusters, apps, builds, security, incidents, cost…), plus the analytics engine — root cause analysis, resource optimization and predictive scaling. Services own their data in PostgreSQL and publish what they learn as events.

4. Kubernetes-as-a-Service​

The clusters KubeOpera runs for you. Each host cluster has a standard control plane; on top of it, every tenant gets one or more KubeSpaces — a virtual cluster (vCluster) with its own API server, backed by dedicated node pools, and with KubeOpera's billing and resource controllers and AI agents running alongside.

5. Platform tools​

Proven open-source tools that KubeOpera configures and drives on your behalf: Flux for delivery, Prometheus/Loki/Jaeger for observability, Trivy and Kyverno for scanning and policy, and cloud workload identity for IAM.

Following a request​

Here is what happens when a developer opens an application's page:

  1. The browser requests /apps/checkout. The Next.js server renders the page and calls /api/apps/checkout.
  2. The API route validates the session JWT and checks the caller may see this tenant's apps.
  3. The route proxies to kubeopera-api, which reads the app's record from PostgreSQL (through the cache service for hot data).
  4. For live status, kubeopera-api asks the tenant's vCluster for the Deployment's current state.
  5. The page streams in health, cost and recent events from k8s-monitor, cost and incident-manager in parallel.

Following an event​

Data moves through the platform on RabbitMQ topic exchanges. The reactive AI pipeline is the most important example: it detects problems, acts, measures the result, and tunes itself — with no human intervention needed for actions you have approved in advance.

Reactive AI Pipeline — Event Flow

Autonomous detection → action → feedback, with feedback tuning detection in a closed loop.

MetricPoint
observability.telemetry
Decisions
analysis.decisions
action.outcomes
feedback.signals
Insights
analysis.insights
Select any service to see its role in the event pipeline.
  1. k8s-monitor polls every registered cluster every 30 seconds and produces a MetricPoint.
  2. observability-agent stores a telemetry snapshot and publishes it to observability.telemetry.
  3. analysis-agent scores every metric against its adaptive threshold and publishes:
    • decisions (analysis.decisions) for the action agent, and
    • insights (analysis.insights) for the recommendation agent and agent runtime.
  4. action-agent executes auto-approved decisions against the Kubernetes API and reports each ActionOutcome.
  5. feedback-agent compares the next snapshot with the pre-action baseline and publishes a reinforcement or correction signal.
  6. analysis-agent consumes those signals and adjusts its thresholds — closing the loop.

In parallel, recommendation-agent turns insights into prioritized recommendations, and agent-runtime automatically launches an SRE investigation when the risk score passes 70.

Read Reactive AI Pipeline for a deeper walkthrough.

Desired state and delivery​

Applications, tenants and clusters are declared as custom resources:

ResourceReconciled byResult
HostClusterhost-cluster-controllerA registered or newly provisioned cluster with Flux and a GitOps tree of its own.
Tenanttenant-controllerAn isolated tenant with its vCluster, quotas and RBAC.
KubeOperaAppapp-controllerGenerated manifests committed to Git and delivered by Flux into the tenant's vCluster.
AdvisorAgentadvisor-controllerA per-app AI advisor that watches the app and recommends improvements.

Because the output is always Git plus Flux, every change is reviewable and revertible, and drift is corrected automatically.

Deployment topology​

KubeOpera itself is deployed with GitOps. An environment is a directory in the deployment repository that Flux reconciles — cluster add-ons (cert-manager, ingress-nginx, monitoring, policy), PostgreSQL, RabbitMQ and every platform service. Platform services run in the kubeopera-* namespaces, with most of them in kubeopera-core.

See Installation to deploy your own environment and Setup for how an environment is structured.

Next steps​

  • Components — what each component is responsible for.
  • Core Services — operate and configure individual services.
  • AI Agents — the autonomous pipeline and reasoning runtime.