Multi-Tenancy
KubeOpera is a multi-tenant platform built on a global ("host") Kubernetes cluster that provisions isolated environments called CloudSpaces for each customer. Every CloudSpace is backed by a vCluster — a fully isolated Kubernetes-in-Kubernetes environment running inside the host cluster.
Two User Roles
All KubeOpera users belong to exactly one of two roles. The role is embedded in the JWT and determines everything: which dashboard is shown, which API data is visible, and which AI tools are available.
| SuperAdmin | Customer | |
|---|---|---|
| Access | Global host cluster | Own CloudSpaces only |
| Dashboard | Cluster-wide monitoring, node pools, incidents, pipelines, analytics, security | CloudSpace overview, app health, AI Advisor |
| AI Chat | All 25 tools (cluster, anomaly, scaling, incidents, …) | App-advisor tools only (get_app_profile, get_app_advice) |
| Assigned by | Manual (seed/admin CLI) | Automatically on signup |
| Redirect after login | /dashboard | /home |
SuperAdmin manages the platform itself: the host cluster, all tenants, infrastructure services, and global agents. The current cluster-monitoring dashboard is the SuperAdmin view.
Customer is the role every new registrant receives. They see only their own CloudSpaces and apps. The AI chat's tool list is filtered down to just get_app_profile/get_app_advice for this role — worth knowing that neither of these two tools currently checks the caller's tenant against the requested app at the data layer; isolation today depends on not already knowing another tenant's exact app identifier, a known gap tracked for a network-level fix (see App Advisor). A third tool name, get_app_live_metrics, appears in the customer allowlist but was never actually implemented — calling it is currently a silent no-op.
CloudSpaces
A CloudSpace is a vCluster provisioned inside the host cluster's vcluster-{name} namespace. It acts as a completely isolated Kubernetes control plane — separate API server, etcd, and scheduler — while sharing the host cluster's worker nodes.
Host Cluster
├── kubeopera-system/ ← platform controllers
│ └── tenant-controller
├── kubeopera-core/ ← platform services
│ └── kubeopera-api, auth-service, k8s-monitor, …
├── vcluster-alice-default/ ← CloudSpace for Alice
│ ├── vCluster control plane
│ └── Alice's apps (auth-service-default, payments-api, …)
└── vcluster-bob-default/ ← CloudSpace for Bob (isolated)
└── Bob's apps
Key properties:
- Full Kubernetes API — every resource a tenant creates (Pods, Services, Ingresses, HPAs) lives inside their CloudSpace and is invisible to other tenants.
- Shared worker nodes — compute is shared across CloudSpaces; resource quotas (
quota.cpu,quota.memory) cap each vCluster's control-plane pod. - Default CloudSpace — every new customer gets exactly one CloudSpace created automatically during onboarding, named
space-{userId[:8]}. - Kubeconfig stored as a Kubernetes Secret
tenant-{cloudSpaceID}-kubeconfiginkubeopera-system.
Tenant Isolation Model
Data isolation
| Layer | How it's isolated |
|---|---|
| Database | tenant_id column on apps table; all queries filtered by tenant_id from JWT |
| API | JWT middleware extracts tenant_id and user_type; Customer requests see only their rows |
| AI Chat | Tool list filtered at request time; system prompt appended with tenant_id context |
| Agent-runtime | Designed to give each vCluster its own app-advisor-srv instance via an AdvisorAgent CR, but nothing currently creates that CR — today there's a single shared app-advisor-srv instance serving every tenant, isolated only by the app-identifier gap noted above |
| Kubernetes | vCluster provides full etcd/API-server isolation; host cluster namespaces are separate |
JWT claims
The auth-service embeds two claims in every token:
| Claim | Type | Value |
|---|---|---|
user_type | string | "super_admin" or "customer" |
tenant_id | string | UUID of the user's default CloudSpace (empty for SuperAdmin) |
user_type is derived from roles at login time — if the user has the "Super Admin" role, user_type is "super_admin"; otherwise "customer". tenant_id is looked up from the database (the CloudSpace with is_default = true and owner_id = user.id).
Provisioning Flow
When a new customer completes email verification:
1. Frontend: POST /api/kubeopera/onboarding/provision
↓
2. kubeopera-api: POST /api/v1/cloudspaces (auth-service)
→ creates CloudSpace record (status: provisioning)
↓
3. kubeopera-api: POST /api/v1/users/{id}/cloudspace (auth-service)
→ marks CloudSpace as default for user
↓
4. kubeopera-api: creates Tenant CR in host cluster
↓
5. tenant-controller reconciles Tenant CR (phase: Pending → Provisioning → InstallingFlux → Ready):
→ helm install vcluster (chart from https://charts.loft.sh)
→ polls readiness
→ extracts kubeconfig from the vCluster's own generated Secret
→ stores it as kubeopera-system/tenant-{csId}-kubeconfig
→ notifies kubeopera-api so the CloudSpace's own DB status moves to active
↓
6. Customer logs in → JWT now carries tenant_id
→ frontend routes to /home (tenant workspace)
Onboarding Experience
After email verification, the customer lands on the /onboarding page. This page triggers provisioning automatically and shows a three-step wizard:
- Provisioning spinner — calls
POST /api/kubeopera/onboarding/provisionon mount - Space confirmation — shows the CloudSpace name once provisioning completes
- Done — redirects to
/home
The customer can use the platform while tenant-controller finishes installing the vCluster in the background. CloudSpace status (provisioning → active) is visible on the /cloudspaces page.
Re-triggering Provisioning
If provisioning fails (network error, Helm timeout), the Tenant CR phase is set to Failed. An operator can re-trigger reconciliation by patching the phase:
kubectl patch tenant alice-default \
--type merge \
-p '{"status":{"phase":"Pending"}}' \
-n kubeopera-system
The controller will restart the Helm install from the beginning.