Skip to main content
Version: 2.0

Cluster Provisioner

Service: cluster-provisioner · Namespace: kubeopera-core · Jobs run in: kubeopera-provisioning

cluster-provisioner turns a Create New request from Cluster Management into a real cluster on your cloud account. It generates a Terraform configuration for the request, runs it as a Kubernetes Job, and hands back a working kubeconfig. host-cluster-controller calls it; you rarely need to.

It's a plain HTTP service rather than an operator: a Terraform apply takes 10 to 40 minutes, which suits a job you start and poll far better than a reconcile loop.

How a provisioning run works​

POST /api/v1/provisions
│ resolve the cloud credential from auth-service (fail fast if it's missing)
│ record the job (status: queued)
▼
Render a Terraform configuration for this cluster (stored as a ConfigMap)
▼
Start Job "provision-<id>" in kubeopera-provisioning
├── init: clone the Terraform modules
└── terraform-runner:
ensure the state bucket exists
terraform init / apply
retrieve the kubeconfig
publish it to Secret "provision-<id>-outputs"
▲
│ polled by host-cluster-controller

Credentials are mounted into the Job from a Secret, never written into the Job spec or Terraform state.

What gets built​

Each run generates a Terraform configuration from the request using KubeOpera's reusable modules:

TopologyModulesResult
AWS vanilla (vanilla_kubeadm)vpc, security, computeA VPC, self-bootstrapping control-plane instances and a worker Auto Scaling Group.
AWS EKS (eks)vpc, eksA VPC, an EKS control plane and a managed node group.
GCP GKE (gke)gcp-network, gkeA VPC network, a GKE control plane and a node pool.

Each cluster gets its own Terraform state bucket (kubeopera-terraform-state-<environment>), never shared.

Vanilla clusters bootstrap themselves​

Vanilla control-plane nodes set themselves up entirely from their instance user data — container runtime, kubeadm init with cloud-provider: external, Calico and the AWS Cloud Controller Manager — with no SSH. They publish their join credentials and admin kubeconfig to SSM Parameter Store for the next step.

With more than one control-plane node, the configuration adds a Network Load Balancer in front of every API server and advertises it as the cluster endpoint. The first node runs kubeadm init --upload-certs; each additional node joins with kubeadm join --control-plane only after the previous one has finished, coordinated through SSM, because concurrent etcd joins can corrupt the cluster. kubeadm handles the etcd membership itself.

Getting the kubeconfig​

TopologyHow
VanillaAsynchronous: the control plane publishes its admin kubeconfig to SSM once it has booted. terraform-runner polls for it (every 20 seconds, up to 20 minutes).
EKS and GKESynchronous: the cluster is ready when apply finishes. terraform-runner builds a kubeconfig from the cluster's endpoint and CA, using short-lived cloud tokens rather than stored credentials.

The kubeconfig is published as a Secret and read back once by host-cluster-controller. cluster-provisioner never stores kubeconfig contents in its database.

Cloud credentials​

cluster-provisioner resolves the chosen credential from auth-service once per run:

  • AWS access key — used directly.
  • AWS assume role — terraform-runner assumes the role (with the external ID, if set) for the run, and EKS kubeconfigs assume it for each token.
  • GCP service account — used for the run and for GKE tokens.

Tearing down​

DELETE /api/v1/provisions/{id} runs terraform destroy for a cluster KubeOpera provisioned, as a Job exactly like provisioning, and removes its state bucket once the destroy succeeds. The dashboard's Delete action on a provisioned cluster uses this.

API​

Only KubeOpera's own services call this API, using their service identities.

MethodPathDescription
POST/api/v1/provisionsStart provisioning. Returns 202 immediately with the job.
GET/api/v1/provisions/{id}Status and message, reconciled with the Job's live state on every read.
GET/api/v1/provisions/{id}/logsJob logs.
GET/api/v1/provisions/{id}/kubeconfigThe kubeconfig, once the job has succeeded (409 before then).
DELETE/api/v1/provisions/{id}Destroy the cluster's infrastructure.
GET/healthzHealth check.

Permissions​

Two identities with deliberately different scope:

# cluster-provisioner: manage Jobs, their inputs, and read their outputs
rules:
- apiGroups: [batch]
resources: [jobs]
verbs: [get, list, watch, create, delete]
- apiGroups: [""]
resources: [pods, pods/log]
verbs: [get, list, watch]
- apiGroups: [""]
resources: [secrets, configmaps]
verbs: [get, list, create, update, delete]
---
# terraform-runner (each Job): only write its own outputs Secret
rules:
- apiGroups: [""]
resources: [secrets]
verbs: [create, update]

Operating it​

kubectl logs -n kubeopera-core deploy/cluster-provisioner -f
kubectl get jobs -n kubeopera-provisioning
kubectl logs -n kubeopera-provisioning job/provision-<id> -f

Configuration​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection.
JOB_NAMESPACEkubeopera-provisioningWhere provisioning Jobs run.
TERRAFORM_RUNNER_IMAGEpinned releaseThe terraform-runner image.
MODULES_REPO_URL—Repository containing the Terraform modules.
AUTH_SERVICE_BASE_URLhttp://authapi:8081For resolving cloud credentials.