Cluster Provisioner
Cluster Provisioner runs real Terraform against a real cloud account on behalf of Host Cluster Controller, turning a fresh-cluster request from the Cluster Setup Wizard into an actual VPC and Kubernetes cluster. It's a plain chi + bun + goose HTTP service like the rest of this platform's backend, deliberately not a controller-runtime operator — a real terraform apply can run 10-40+ minutes, and a reconcile loop's cheap-per-pass assumptions and requeue/backoff machinery are the wrong shape for that. host-cluster-controller talks to it purely over HTTP, the same trusted-internal-caller pattern build-service/cicd-gateway already use elsewhere.
Architecture: one Job per provisioning run
Modeled directly on build-service's own Kaniko-per-build pattern — one Kubernetes Job per operation, in a dedicated namespace (kubeopera-provisioning), credentials mounted from a Secret never inlined into the Job spec.
POST /api/v1/provisions
│
│ resolve credential from auth-service (once, fail fast)
│ persist a provision_jobs row (status=queued)
▼
ensureEnvConfigMap() renders a real Terraform root module
│ (main.tf + backend.tf) for this one job,
│ stored as a ConfigMap
▼
StartJob() creates batchv1.Job "provision-<id>" in
│ kubeopera-provisioning:
│ initContainer: git-clone gitops-iac
│ (for its existing modules/)
│ container "terraform": the
│ terraform-runner image, ConfigMap
│ volume-mounted over the one new
│ environments/<name>/ directory the
│ clone doesn't already have
▼
terraform-runner (inside the Job):
ensureBackendBucket() idempotent S3 state bucket create
terraform init / apply via terraform-exec (real init/apply
calls, not parsed CLI stdout)
kubeconfig retrieval branches on topology -- see below
publish outputs Secret "provision-<id>-outputs", written via
its own narrowly-scoped ServiceAccount
▲
│ polled by GetStatus / fetched once by GetKubeconfig
│
host-cluster-controller
Generated Terraform environments
Nothing commits a per-cluster environment directory to gitops-iac ahead of time — RenderTerraformEnvironment generates one on demand from the wizard's inputs, wiring the existing, unmodified Terraform modules:
-
vanilla_kubeadm—vpc+security+computemodules, withcompute'senable_self_bootstrap_control_plane = true. That flag swaps the control-plane node'suser_datafrom a trivial placeholder to a real self-join script: full OS/containerd/kubelet setup,kubeadm initwith the platform's own hard-woncloud-provider: externalfix baked in, Calico CNI, the AWS Cloud Controller Manager, then publishing its own join credentials and admin kubeconfig to SSM — mirroring the worker ASG's own already-proven self-bootstrapuser_data, extended to the control plane. It's opt-in and additive: every hand-maintained environment (development/staging/production) leaves the flag unset and keeps its original bastion-driven bootstrap unchanged. Provisioning also creates theaws_ssm_parameterresources (join_token,ca_cert_hash,admin_kubeconfig, all seededPENDING) that node's own boot process later overwrites.ControlPlaneCount > 1adds a real Network Load Balancer (aws_lb.k8s_api+ target group + listener, rendered inline in the generated environment — deliberately notmodules/loadbalancer, which pulls in ACM/ALB assumptions this doesn't need) fronting every control-plane node's API server on port 6443, and switchescontrol_plane_endpoint_override(amodules/computevariable, validated as required whenevercontrol_plane_count > 1) to the NLB's own DNS name instead of defaulting to node 0's private IP — otherwise a multi-node cluster's advertised API endpoint would point at one node and defeat the point of having more than one. Theadmin_kubeconfig/k8s_api_endpointSSM parameters correspondingly point at the NLB's DNS name rather than a single node's IP once this is set. A newaws_ssm_parameter.k8s_certificate_key(also seededPENDING) carries the--certificate-keynode 0 generates forkubeadm init --upload-certs, which every additional control-plane node needs tokubeadm join --control-planewith its own copy of the control-plane certificates. Join ordering across control-plane nodes is enforced with a strict SSM-parameter relay (control_plane_joined/<index>, each node waiting for the previous index before joining) — concurrent etcd joins can corrupt the cluster's etcd state, so this is deliberately serialized rather than left to chance; kubeadm itself still owns the actual etcd learner-add/promotion mechanics. -
eks—vpc+ the neweksmodule (aws_eks_cluster+ a managedaws_eks_node_group+ their IAM roles). No SSM parameters at all — nothing asynchronous on this path.
Both templates deliberately stay minimal: no loadbalancer/route53/ACM/backup-bucket/cost-explorer wiring, even though environments/development's own hand-written main.tf has all of that — those are the team's own platform concerns, not appropriate defaults for a freshly self-service-provisioned cluster. The backend state bucket follows the existing per-environment convention exactly: kubeopera-terraform-state-<TerraformEnvironment>, S3-native locking, never a shared bucket.
Kubeconfig retrieval: two different strategies
This is the one place the two topologies genuinely diverge, because the two clusters become reachable at different points relative to terraform apply finishing:
vanilla_kubeadm— asynchronous. The control-plane node runskubeadm init+ Calico + the AWS Cloud Controller Manager entirely from its own EC2user_data, well after Terraform itself has finished creating the instance, then publishes its ownadmin.conf(base64, as aSecureString) to theadmin_kubeconfigSSM parameter the generated environment already created.terraform-runnerpolls that parameter (every 20s, up to 20 minutes) rather than reading a same-applyTerraform output, which would only ever see the still-PENDINGplaceholder.eks— synchronous. Bothaws_eks_clusterandaws_eks_node_groupblock withinterraform applyitself until they're genuinely ready, soterraform-runnerreads the cluster'scluster_name/cluster_endpoint/cluster_ca_certificateoutputs immediately afterward and constructs a real kubeconfig in Go (a proper YAML marshal, not string formatting) — no polling needed. The AWS credential itsexec:auth plugin needs is embedded directly into the kubeconfig at construction time, deliberately not round-tripped through a Terraform variable or output —terraform-runneralready holds it in-process from resolving it for the provider itself, and state files are a well-known place a credential can leak if it doesn't need to be there.
Either way, the resulting kubeconfig is published as a Secret in kubeopera-provisioning and read back by cluster-provisioner's own GetKubeconfig — never cached in Postgres. The provision_jobs table only ever stores that Secret's (deterministic, computable) name, not its content.
Cloud credentials
Provisioning needs a real AWS credential, resolved exactly once per job, immediately before the Job is created — see auth-service's cloud_credentials store (platform-admin-only, no tenant dimension at all, since provisioning a host cluster is inherently a platform operation). Only access_key credentials work end-to-end today; an assume_role credential can be stored but nothing yet performs the actual sts:AssumeRole call a real provisioning run or an EKS kubeconfig's exec plugin would need to use one.
API
All routes below /api/v1 are behind a shared X-Api-Key — trusted-caller-only, the same convention as build-service.
| Method | Path | Purpose |
|---|---|---|
POST | /api/v1/provisions | Start a provisioning run. Returns 202 with the new job immediately — never blocks on Terraform itself. |
GET | /api/v1/provisions/{id} | Poll status/message. Sync-on-read: reconciles the stored row against the Job's live k8s state on every call, the same pattern AppService.getAppSynced already established. |
GET | /api/v1/provisions/{id}/logs | Real Job pod logs — a plain synchronous full-text fetch, not a stream, matching build-service's own identical choice for the same reason (no SSE/streaming transport exists elsewhere in this platform for this class of operation). |
GET | /api/v1/provisions/{id}/kubeconfig | Fetch the real kubeconfig content, once the job has succeeded. 409 if it hasn't. |
RBAC
Two separate ServiceAccounts, deliberately scoped differently:
# cluster-provisioner itself: create/poll/clean up Jobs, its own
# per-job credential Secret, and the per-job generated-environment
# ConfigMap. Never writes the *outputs* Secret -- only reads it back.
rules:
- apiGroups: [batch]
resources: [jobs]
verbs: [get, list, watch, create, delete]
- apiGroups: [""]
resources: [pods]
verbs: [get, list, watch]
- apiGroups: [""]
resources: [pods/log]
verbs: [get]
- apiGroups: [""]
resources: [secrets]
verbs: [get, list, create, update, delete]
- apiGroups: [""]
resources: [configmaps]
verbs: [get, list, create, update, delete]
---
# terraform-runner (every Job runs as this identity): only ever
# creates/updates its own job's outputs Secret by name -- no reason to
# ever read another job's credential Secret or anything else here.
rules:
- apiGroups: [""]
resources: [secrets]
verbs: [create, update]
Deployment
Single-replica Deployment in kubeopera-core. The terraform-runner image it references is a separate service in the same monorepo — pinned to an exact tag by hand in terraform_job.go (no gitops-managed Deployment ever runs that image directly, so there's nothing for an automated GitOps update to bump; it's referenced only from this one Go constant, the same way build-service hand-pins the upstream Kaniko executor image it shells out to).
kubectl logs -n kubeopera-core deploy/cluster-provisioner -f
kubectl get jobs -n kubeopera-provisioning
kubectl logs -n kubeopera-provisioning job/provision-<id> -f
Health check at GET /health.