Skip to main content
Version: 1.0

Clusters

/clusters is Cluster Management — where a platform admin links KubeOpera to the Kubernetes cluster(s) it runs workloads on. An admin can register the cluster KubeOpera is already running in, connect an existing cluster by kubeconfig, or have KubeOpera provision a brand-new cluster on AWS from scratch.

/clusters/multi is a separate page — the fleet-wide comparison view — covered at the bottom of this page.

The core model: clusters + HostCluster​

Every cluster KubeOpera knows about is a row in kubeopera-api's own clusters table — id, name, provider, region, source, is_host, status, message, kubernetes_version, node_count/ready_node_count, timestamps. source records how the row came to exist: self_registered, connected_existing, provisioned_vanilla, or provisioned_eks. At most one row can ever have is_host = true — enforced with a partial unique index at the database layer, not just application logic — since KubeOpera itself only ever runs its own control plane on one cluster at a time.

For every row except a self-registered one, there's a matching HostCluster custom resource (group kubeopera.io/v1alpha1, always in the host cluster's own kubeopera-system namespace) that a dedicated controller — host-cluster-controller — reconciles. The clusters table row is what the UI reads; the HostCluster CR is what's actually doing the work. kubeopera-api keeps the two in sync the same way it already does for apps (AppService.getAppSynced): every GET /clusters/GET /clusters/{id} call re-reads the CR's live phase and message and writes back any change, rather than running a separate background poller.

apiVersion: kubeopera.io/v1alpha1
kind: HostCluster
metadata:
name: <clusters.id> # the CR's own name is the UUID, no slugify needed
namespace: kubeopera-system
spec:
clusterID: "7e9d3a1b-..."
mode: connect_existing # | provision_vanilla | provision_eks
provider: aws
region: eu-north-1
connectKubeconfigSecretRef: hostcluster-7e9d3a1b-input-kubeconfig # connect_existing only
provisioning: # provision_vanilla / provision_eks only
topology: vanilla_kubeadm
controlPlaneCount: 1 # vanilla_kubeadm only; >1 provisions an NLB
nodeCount: 2
instanceType: t3.medium
kubernetesVersion: "1.31.0"
credentialRef: <cloud_credentials.id>
terraformEnvironment: cluster-7e9d3a1b-...
gitOpsRepoURL: https://github.com/ochestra-tech/gitops-iac.git
gitOpsPath: fluxcd/clusters/my-cluster
gitCredentialsSecretRef: gitops-iac-credentials
domain: my-cluster.acme.com
registryPullSecretRef: hostcluster-7e9d3a1b-registry-pull # optional -- a customer-supplied registry credential, in place of the platform default
adminEmail: admin@acme.com # optional -- omit both to skip automatic admin-account creation
adminUsername: admin
status:
phase: Ready # Pending | ValidatingAccess | Provisioning |
# BootstrappingFlux | GeneratingGitOpsTree |
# InstallingKubeOpera | Ready | Failed | Deleting
kubeconfigSecret: hostcluster-7e9d3a1b-kubeconfig
fluxHealthy: true
adminPasswordSecretRef: hostcluster-7e9d3a1b-admin-password
message: "flux healthy, pointed at fluxcd/clusters/my-cluster. super-admin account admin@acme.com created -- retrieve the password from secret kubeopera-system/hostcluster-7e9d3a1b-admin-password"

Whichever of the three onboarding paths below is used, they all converge on the same last three phases: BootstrappingFlux (installs Flux via its own Helm chart onto the target, waits for it to report healthy, seeds it a git-credentials Secret), GeneratingGitOpsTree (calls gitops-scaffolder to generate and push a complete, customer-specific tree at gitOpsPath — cert-manager, ingress-nginx, storage, database, and every service, each with this cluster's own domain — before anything points Flux at it; see the Setup Guide for exactly what gets generated), and InstallingKubeOpera (create a GitRepository + Kustomization pointed at that now-real tree). Once apps itself reports Ready, and only if adminEmail/adminUsername were set, a one-time call creates the initial super-admin account before the CR moves to Ready — status.message says plainly whether that happened, and status.adminPasswordSecretRef names where the generated password landed. Ready still doesn't, by itself, confirm every individual service is healthy — only that Flux is pointed at a real, complete tree and the platform's own apps layer has converged.

Getting started: the empty state​

A KubeOpera deployment with no is_host = true row shows a Get Started panel instead of the (empty) cluster list — deliberately triggered by "no host cluster," not "no clusters at all," so a deployment that has only connected non-host clusters still prompts for a real host. There's also a dismissible banner on the main dashboard pointing here until a host cluster exists.

  • Register This Cluster — the fast path, covered below.
  • Set Up Manually — opens the same Cluster Setup Wizard the "New Cluster" button opens.

Self-registration​

If you're looking at this page from a KubeOpera instance that's already running somewhere — which is the common case, since KubeOpera itself has to be deployed via the Setup Guide before there's a dashboard to open at all — Register This Cluster is one click. kubeopera-api inspects the cluster it's already running in using its own in-cluster ServiceAccount: no kubeconfig, no cloud credential, nothing to type in.

  • Node count / readiness — a straightforward kubectl get nodes from kubeopera-api's own pod.
  • Kubernetes version — Discovery().ServerVersion().
  • Provider / region — read from any node's .spec.providerID. AWS's Cloud Controller Manager sets this to aws:///<az>/<instance-id>, so a prefix match gives the provider and the AZ segment (minus its trailing letter) gives the region. A cluster with no CCM installed — a bare kubeadm cluster with no cloud integration — reports provider: unknown rather than guessing.
  • Idempotency — the inspector also computes a stable fingerprint from the target's own kube-system namespace UID (cluster_uid). Clicking the button again returns the existing row rather than erroring or creating a duplicate.

Self-registration never creates a HostCluster CR — there's nothing to reconcile, since the cluster is already running KubeOpera by definition.

Connect Existing​

For a cluster you already have a working kubeconfig for — separately provisioned infrastructure, an existing production cluster, anything KubeOpera didn't create itself. Paste or upload the kubeconfig, pick a name/provider/region, submit.

kubeopera-api does the minimum validation it can up front — confirming the pasted text is at least valid YAML — before creating anything, so a garbage paste fails immediately with a specific message rather than a generic error. The real test, whether the cluster is actually reachable, happens asynchronously once the HostCluster CR exists:

  1. ValidatingAccess — host-cluster-controller builds a client from the stored kubeconfig and calls Discovery().ServerVersion(). A transient failure (a momentary network blip) is retried with backoff for up to 5 minutes before the CR moves to Failed with the real underlying error attached — not a generic "couldn't connect."
  2. BootstrappingFlux → GeneratingGitOpsTree → InstallingKubeOpera → Ready, as described above.

The kubeconfig is stored close to as-supplied — unlike a tenant's vCluster kubeconfig (which KubeOpera rewrites to an in-cluster short-DNS form, since a vCluster is always reachable from inside the host cluster's own network), a connected external cluster has no such guarantee, so its server: field is left exactly as the admin's own cluster exposes it.

Disconnecting a connected cluster removes the HostCluster CR (which triggers a finalizer-driven cleanup of the GitRepository/Kustomization CRs and the stored kubeconfig Secret) but deliberately does not touch anything actually running on the target — it may still be in active use outside KubeOpera. A destructive "fully tear down" action doesn't exist yet.

Create New: provisioning a cluster from scratch​

The newest path, and the one that actually stands up real cloud infrastructure rather than pointing at something that already exists. Today this supports AWS only — GCP and Azure show as visibly disabled "Soon" options in the wizard rather than being hidden, so the choice is already in the right place once a real implementation exists for them.

Choosing a topology​

  • Vanilla Kubernetes — a self-managed kubeadm cluster: KubeOpera provisions the VPC, one or more control-plane EC2 instances (stacked etcd — each control-plane node runs its own etcd member, matching the platform's real dev topology, rather than a separate etcd cluster), and a worker Auto Scaling Group. Every control-plane node bootstraps itself entirely from its own EC2 user_data with no SSH involved at any point — the first node runs kubeadm init --upload-certs, and each additional one runs kubeadm join --control-plane once the previous one has finished (a strict SSM-parameter relay enforces this ordering, since two control-plane nodes joining etcd concurrently can corrupt it; kubeadm itself handles the actual etcd learner-add/promotion, no hand-rolled quorum logic involved). A single-node cluster is still the default and fully supported — asking for more than one just adds a Network Load Balancer in front of the API servers (control_plane_endpoint_override, required whenever more than one control-plane node is requested) so no client ends up pointed at one specific node's own IP.
  • Managed (EKS) — AWS operates the control plane; KubeOpera provisions the VPC and a managed EKS node group. No bootstrap script at all on this path — EKS's own managed node group mechanism handles kubelet/containerd installation and cluster join declaratively.

Cloud credentials​

Provisioning needs real AWS credentials — an access key/secret pair, stored via a new, platform-admin-only credential store in auth-service (encrypted at rest, the plaintext never round-tripped back to the UI after creation; a stored credential is shown afterward only as a label and the last 4 characters). The Cluster Setup Wizard's credential step lets you pick an existing one or add a new one inline. Only access_key credentials work end-to-end today; assume_role credentials can be stored but nothing yet performs the actual role-assumption needed to use one for a real provisioning run.

What actually runs​

Submitting the wizard doesn't run Terraform inline — kubeopera-api creates the clusters row and the HostCluster CR (mode: provision_vanilla or provision_eks) and returns immediately. From there:

  1. Provisioning — host-cluster-controller calls a dedicated service, cluster-provisioner, to start the actual work, and polls it for status on every subsequent reconcile (tracked via an annotation on the CR, so this survives a controller pod restart mid-provision — a real Terraform apply can run 10-40+ minutes).
  2. cluster-provisioner resolves the chosen credential, generates a real Terraform root module wiring the existing vpc/security/compute modules (vanilla) or vpc/eks modules (managed) with the wizard's inputs, and runs it as a dedicated Kubernetes Job — one Job per provisioning request, mirroring how this platform already runs one-Job-per-build for container image builds. terraform init/apply run for real inside that Job.
  3. Once infrastructure exists, cluster-provisioner retrieves a working kubeconfig — the mechanism differs by topology. A vanilla cluster's control plane publishes its own admin kubeconfig to AWS SSM Parameter Store once it's actually finished booting (asynchronous — this is why the retrieval step polls rather than reading a Terraform output directly). An EKS cluster's kubeconfig is available synchronously the moment apply finishes (both the cluster and its node group block within apply until they're genuinely ready), constructed directly from the cluster's endpoint and CA certificate, with a short-lived AWS-signed token for authentication rather than a long-lived credential.
  4. host-cluster-controller fetches that kubeconfig once, stores it the same way a Connect-Existing kubeconfig is stored, and the CR falls through into the identical BootstrappingFlux → GeneratingGitOpsTree → InstallingKubeOpera → Ready sequence described above — provisioning a cluster and connecting an existing one converge into the same code path the moment a working kubeconfig exists.

Watching progress​

The wizard itself closes immediately after submitting — deliberately. An earlier version of the (unrelated) application deploy wizard stayed open through its own long-running operation and got mistaken for a stuck dialog, prompting repeated resubmission; a cluster provisioning run is far longer than that ever was, so the same mistake would be worse here. Progress instead shows on the cluster's own row in the list — status badge plus the CR's own message field — which kubeopera-api keeps in sync on every page load the same sync-on-read way described above.

For a lower-level view than the dashboard shows, the actual terraform apply output is real Job logs:

kubectl get hostcluster -n kubeopera-system -w
kubectl get jobs -n kubeopera-provisioning
kubectl logs -n kubeopera-provisioning job/provision-<id> -f

All Clusters (/clusters)​

Displays every row from the real clusters table as a card — health badge, node count, provider, region, Kubernetes version, and (for a self-registered cluster only) a per-node breakdown. A card's status badge reflects either the clusters table's own status (self-registered) or the synced HostCluster phase (everything else) — pending/provisioning/validatingaccess/bootstrappingflux/installingkubeopera all render as an in-progress amber state, ready/running/active as green, failed/error as red.

Cluster Health Score​

For a cluster with live metrics available, the badge comes from k8s-monitor's real health-scoring logic — this part predates Cluster Management and is unchanged by it. It starts at 100 and subtracts fixed penalties per issue found, rather than combining weighted percentage sub-scores:

ConditionPenalty
Each not-ready node−5
Each node under memory pressure−3
Each node under disk pressure−4
Each node under PID pressure−2
Each node with network unavailable−6
Each failed pod−2
Each crash-looping pod−2.5
More than 10 restarting pods−5 (flat)
Control plane not fully healthy−30 (or −15 if only the API server is affected)
Each unhealthy control-plane component−8
Cluster CPU ≥ 95% (≥ 80%: −5)−10
Cluster memory ≥ 95% (≥ 80%: −5)−10

The result is clamped to 0–100. A badge of green/amber/red follows the usual ≥80 / 60–79 / below-60 bands.

Multi-Cluster View (/clusters/multi)​

This page calls kubeopera-api's federation-overview endpoint, which fans out concurrently to every cluster registered in a CLUSTER_REGISTRY configuration value — real code, genuinely built for a multi-cluster fleet, and unrelated to the clusters table Cluster Management now manages (this is a separate, older, config-driven registry, not yet wired to read from the same table). In a typical deployment CLUSTER_REGISTRY isn't set, so it falls back to a single entry representing the host cluster itself — the fan-out logic is real, but there's only one cluster to fan out to until a second one is actually registered there too.

Summary row — four real metric cards: Total Clusters, Healthy Clusters (ready/total), Total Nodes (ready/total), and Avg CPU Usage. There's no fleet-wide weighted health score or a monthly-cost figure shown on this page.

Comparison table — sortable, showing each cluster's health, node readiness, and CPU/memory usage bars.

Treemap — a recharts Treemap, sized specifically by CPU usage percentage (not a general "resource consumption" blend of CPU and memory), colored by cluster status. Hovering shows the cluster name and its CPU percentage.

Cost data does exist and is real — k8s-monitor's cost endpoint computes it from live node/pod resource data against a real pricing table, and kubeopera-api's federation-overview response includes an extrapolated monthly figure derived from it — but this page doesn't currently render that figure anywhere in its UI.

App Lifecycle​

When a user creates an application via the UI, kubeopera-api:

  1. Persists the app record to PostgreSQL
  2. Creates a KubeOperaApp Kubernetes custom resource

The app-controller then reconciles the CRD:

  1. Generates Kubernetes manifests via app-service
  2. Pushes manifest files to the fleet Git repository
  3. Creates or updates a Flux GitRepository and Kustomization CR — inside the target tenant's own vCluster, not the host cluster
  4. Updates the CR status to Ready once Flux reconciles

See App Creation Flow for the full, real deploy sequence, and KubeOpera API for the app endpoints themselves.