Architecture
KubeOpera is made of two codebases that work together: a Next.js dashboard and a set of Go microservices running in Kubernetes. They communicate in three ways:
- REST — for requests a person or another service makes and waits for.
- RabbitMQ topic exchanges — for events that many services react to independently.
- Kubernetes custom resources — for desired state that controllers reconcile.
This page walks through the system layer by layer, then follows a request and an event through it.
System diagram
Select any component to see what it does.
High-Level System Architecture
Select any component to see what it doesThe layers
1. User interface
Everyone — platform engineers, developers, SREs and FinOps — works in the same dashboard. Pages are server-rendered where possible and stream live data (agent runs, logs, health) where it matters.
2. API gateway
The dashboard never calls a backend service directly from the browser. Every call goes to a Next.js route under /api/*, which:
- Authenticates the request using the session's JWT issued by auth-service.
- Rate-limits per endpoint to protect downstream services.
- Enforces an allowlist — only known service hostnames can be proxied to, which prevents server-side request forgery.
- Routes the request to the right service based on its path prefix.
- Load-balances across service replicas through Kubernetes Services.
Keeping backend URLs and credentials on the server also means no CORS configuration and no secrets in the browser.
3. Central management plane
This is the heart of the platform: Go services, each owning one domain (clusters, apps, builds, security, incidents, cost…), plus the analytics engine — root cause analysis, resource optimization and predictive scaling. Services own their data in PostgreSQL and publish what they learn as events.
4. Kubernetes-as-a-Service
The clusters KubeOpera runs for you. Each host cluster has a standard control plane; on top of it, every tenant gets one or more KubeSpaces — a virtual cluster (vCluster) with its own API server, backed by dedicated node pools, and with KubeOpera's billing and resource controllers and AI agents running alongside.
5. Platform tools
Proven open-source tools that KubeOpera configures and drives on your behalf: Flux for delivery, Prometheus/Loki/Jaeger for observability, Trivy and Kyverno for scanning and policy, and cloud workload identity for IAM.
Following a request
Here is what happens when a developer opens an application's page:
- The browser requests
/apps/checkout. The Next.js server renders the page and calls/api/apps/checkout. - The API route validates the session JWT and checks the caller may see this tenant's apps.
- The route proxies to kubeopera-api, which reads the app's record from PostgreSQL (through the cache service for hot data).
- For live status, kubeopera-api asks the tenant's vCluster for the Deployment's current state.
- The page streams in health, cost and recent events from k8s-monitor, cost and incident-manager in parallel.
Following an event
Data moves through the platform on RabbitMQ topic exchanges. The reactive AI pipeline is the most important example: it detects problems, acts, measures the result, and tunes itself — with no human intervention needed for actions you have approved in advance.
Reactive AI Pipeline — Event Flow
Autonomous detection → action → feedback, with feedback tuning detection in a closed loop.
- k8s-monitor polls every registered cluster every 30 seconds and produces a
MetricPoint. - observability-agent stores a telemetry snapshot and publishes it to
observability.telemetry. - analysis-agent scores every metric against its adaptive threshold and publishes:
- decisions (
analysis.decisions) for the action agent, and - insights (
analysis.insights) for the recommendation agent and agent runtime.
- decisions (
- action-agent executes auto-approved decisions against the Kubernetes API and reports each
ActionOutcome. - feedback-agent compares the next snapshot with the pre-action baseline and publishes a reinforcement or correction signal.
- analysis-agent consumes those signals and adjusts its thresholds — closing the loop.
In parallel, recommendation-agent turns insights into prioritized recommendations, and agent-runtime automatically launches an SRE investigation when the risk score passes 70.
Read Reactive AI Pipeline for a deeper walkthrough.
Desired state and delivery
Applications, tenants and clusters are declared as custom resources:
| Resource | Reconciled by | Result |
|---|---|---|
HostCluster | host-cluster-controller | A registered or newly provisioned cluster with Flux and a GitOps tree of its own. |
Tenant | tenant-controller | An isolated tenant with its vCluster, quotas and RBAC. |
KubeOperaApp | app-controller | Generated manifests committed to Git and delivered by Flux into the tenant's vCluster. |
AdvisorAgent | advisor-controller | A per-app AI advisor that watches the app and recommends improvements. |
Because the output is always Git plus Flux, every change is reviewable and revertible, and drift is corrected automatically.
Deployment topology
KubeOpera itself is deployed with GitOps. An environment is a directory in the deployment repository that Flux reconciles — cluster add-ons (cert-manager, ingress-nginx, monitoring, policy), PostgreSQL, RabbitMQ and every platform service. Platform services run in the kubeopera-* namespaces, with most of them in kubeopera-core.
See Installation to deploy your own environment and Setup for how an environment is structured.
Next steps
- Components — what each component is responsible for.
- Core Services — operate and configure individual services.
- AI Agents — the autonomous pipeline and reasoning runtime.