Skip to main content
Version: 1.0

Tenant Controller

Tenant Controller is a controller-runtime operator that watches Tenant CRs and provisions isolated vCluster environments for each KubeOpera customer. It is the only component that interacts with the Helm SDK and the vCluster lifecycle.

CRD: Tenant​

Group/Version: kubeopera.io/v1alpha1 · Kind: Tenant · Short name: kt

apiVersion: kubeopera.io/v1alpha1
kind: Tenant
metadata:
name: alice-default
namespace: kubeopera-system
spec:
userID: "3f1a2b4c-..." # UUID of the owning user
tenantName: "alice-default" # slug → vcluster release + namespace names
cloudSpaceID: "7e9d3a1b-..." # CloudSpace UUID in auth-service / kubeopera-api
quota:
cpu: "2" # vCluster control-plane pod limit
memory: "4Gi"
status:
phase: Ready # Pending | Provisioning | Ready | Failed
vclusterName: vc-alice-default
namespace: vcluster-alice-default
kubeconfigSecret: tenant-7e9d3a1b-kubeconfig
lastSyncedAt: "2026-04-27T14:05:00Z"

Naming rules:

  • vCluster Helm release name: vc-{tenantName} (max 53 chars)
  • Host namespace: vcluster-{tenantName}
  • Kubeconfig Secret: tenant-{cloudSpaceID}-kubeconfig in kubeopera-system

tenantName must match ^[a-z0-9][a-z0-9-]{1,30}[a-z0-9]$.

Reconcile State Machine​

Tenant CR created / updated
│
┌─────▼──────┐
│ Pending │ (initial state or manual reset)
└─────┬──────┘
│ create namespace vcluster-{name}
│ helm.InstallOrUpgrade(vc-{name}, …)
┌─────▼──────────┐
│ Provisioning │ ← requeues every 15 s
└─────┬──────────┘
│ helm.IsReady() → release.StatusDeployed?
│ GetVClusterKubeconfig() → Secret vc-{name}.config populated?
│ StoreKubeconfig() → kubeopera-system/tenant-{csId}-kubeconfig
│ HTTP PATCH kubeopera-api /api/v1/cloudspaces/{id}
┌─────▼──────┐
│ Ready │ (idempotent — further reconciles no-op)
└────────────┘
│ (on error at any step)
┌─────▼──────┐
│ Failed │ ← operator patches phase to Pending to retry
└────────────┘

Wait = false in Helm install — the reconciler does not block during the vCluster boot sequence. Instead it polls helm.IsReady() every 15 seconds until the Helm release reaches StatusDeployed.

Helm Integration​

The controller uses helm.sh/helm/v3 v3.16.4 Go SDK (helm.sh/helm/v3/pkg/action).

Chart: oci://ghcr.io/loft-sh/charts/vcluster @ 0.21.2

Default Helm values applied to every CloudSpace:

controlPlane:
distro:
k3s:
enabled: true
statefulSet:
resources:
requests: { cpu: "{quota.cpu}", memory: "{quota.memory}" }
limits: { cpu: "{quota.cpu}", memory: "{quota.memory}" }
sync:
toHost:
ingresses: { enabled: true }
serviceAccounts: { enabled: true }
exportKubeConfig:
enabled: true # creates Secret vc-{releaseName} with key "config"

The exportKubeConfig flag causes vCluster to write the kubeconfig for the virtual cluster into a Kubernetes Secret (vc-{releaseName}) in the host namespace, which the controller reads once provisioning is complete.

In-Cluster Helm Configuration​

Helm SDK requires a RESTClientGetter. Inside a controller pod there are no kubeconfig files, so the controller provides a custom restClientGetter that wraps the controller manager's *rest.Config:

// internal/helm/rest_client_getter.go
type restClientGetter struct {
config *rest.Config // in-cluster REST config from ctrl.GetConfigOrDie()
namespace string
}
// Implements: ToRESTConfig, ToDiscoveryClient (memory-cached), ToRESTMapper, ToRawKubeConfigLoader

This allows Helm's install/upgrade/status/uninstall actions to work identically inside the cluster and in local development (with KUBECONFIG set).

Kubeconfig Storage​

After the vCluster is ready:

  1. Reads Secret vc-{releaseName} from the vCluster namespace (key: config).
  2. Creates/updates Secret tenant-{cloudSpaceID}-kubeconfig in kubeopera-system.
# Retrieve a CloudSpace kubeconfig as SuperAdmin
kubectl get secret tenant-<cs-id>-kubeconfig \
-n kubeopera-system \
-o jsonpath='{.data.kubeconfig}' | base64 -d

The Secret is labelled kubeopera.io/tenant-id={cloudSpaceID} and kubeopera.io/managed-by=tenant-controller.

kubeopera-api Callback​

Once the kubeconfig is stored, the controller sends a PATCH to kubeopera-api:

PATCH /api/v1/cloudspaces/{cloudSpaceID}
Content-Type: application/json

{
"vcluster_id": "vc-alice-default",
"namespace": "vcluster-alice-default",
"kubeconfig_secret": "tenant-7e9d3a1b-kubeconfig",
"status": "active"
}

This is non-fatal — if kubeopera-api is unreachable, the controller still sets the Tenant phase to Ready and logs a warning. The CloudSpace status will sync on the customer's next login when the JWT is refreshed.

RBAC​

The tenant-controller ServiceAccount requires broad cluster-level permissions because Helm queries all API groups during install:

rules:
- apiGroups: [kubeopera.io]
resources: [tenants, tenants/status]
verbs: [get, list, watch, update, patch]
- apiGroups: [""]
resources: [namespaces, secrets, configmaps, events]
verbs: [get, list, watch, create, update, patch, delete]
- apiGroups: ["*"]
resources: ["*"]
verbs: [get, list] # Helm discovery only

The get/list on * is read-only — the controller does not write to arbitrary resource types.

Environment Variables​

VariableDefaultDescription
KUBEOPERA_API_BASE_URLhttp://kubeopera-api:8080kubeopera-api endpoint for CloudSpace status callback
INTERNAL_SERVICE_TOKEN(empty)Bearer token for kubeopera-api requests (optional)
ENVIRONMENTdevelopmentPassed through to the Helm values for labelling

Deployment​

The controller runs as a single replica Deployment in kubeopera-system. It mounts an emptyDir volume at /root/.cache/helm for Helm chart caching between reconcile loops, avoiding re-pulling the vCluster OCI chart on every reconciliation.

# Apply CRD (required before controller starts)
kubectl apply -f config/crd/kubeopera.io_tenants.yaml

# Apply RBAC + Deployment
kubectl apply -f deployments/base/deployment.yaml

# Check controller logs
kubectl logs -n kubeopera-system -l app=tenant-controller -f

Re-triggering a Failed Provision​

If a Tenant CR is stuck in Failed, patch the phase to Pending:

kubectl patch tenant alice-default \
--type=merge \
-p '{"status":{"phase":"Pending"}}' \
-n kubeopera-system

The controller will restart the Helm install from scratch. If the Helm release already exists (partial install), InstallOrUpgrade will detect it via action.NewHistory and call Upgrade instead of Install.

Monitoring​

The controller exposes Prometheus metrics on :8080/metrics via controller-runtime's built-in metric server. Key metrics:

  • controller_runtime_reconcile_total — reconcile attempts per controller
  • controller_runtime_reconcile_errors_total — failed reconciles
  • controller_runtime_reconcile_time_seconds — reconcile duration histogram

Health and readiness probes on :8081/healthz and :8081/readyz.