KubeOpera API
Repo: o-apps/kubeopera-api · Port: 8090 · DB Schema: shared kubeopera database, own tables
KubeOpera API is the platform's central orchestration service. It is the only backend that tenants and the frontend talk to for anything involving an application, a CloudSpace, a vCluster, a CI/CD pipeline, or a Kaniko-built image — everything else in the platform (auth, billing, the agentic AI layer) is reached through its own service, but application lifecycle and GitOps deployment all flow through here.
The design decision that shapes almost everything else in this service is that an application's desired state lives in Git, not in a live cluster mutation. Creating, updating, or rolling back an app never talks to Kubernetes directly. Instead, KubeOpera API writes a KubeOperaApp custom resource, a separate controller (app-controller, documented on its own page) reconciles that CR into a generated Kubernetes manifest and pushes it to a Git repository, and Flux — running inside the tenant's own vCluster — is what actually applies it. This indirection is why the API can offer things like deployment history and rollback almost for free: Kubernetes already keeps a full revision history of every meaningful change via ReplicaSets, so KubeOperaAPI reads that history back rather than maintaining its own.
How app deployment actually works
A POST /api/v1/apps call is a single, atomic request — it carries the image, port, TLS/domain preference, resource requests, autoscaling config, environment variables, and health checks all at once, precisely so that no caller ever needs to create an app and then immediately follow up with a config update before the app is in a stable state. Internally, AppService validates the request and persists the row to Postgres first. If APP_CONTROLLER_ENABLED=true (the case in every real deployment; disabled only for local development without a KubeOperaApp CRD installed), it then builds a KubeOperaApp CR from the request — computing a server-side subdomain of the form <app>-<cloudspace-slug>.apps.kubeopera.io rather than trusting whatever the client sent, and setting ingress.tls/the cert-manager.io/cluster-issuer annotation when TLS is requested — and creates that CR against the Kubernetes API. If a caller explicitly targets one of several vClusters within a CloudSpace via tenant_cr_name, the CR carries that identifier so app-controller resolves the named vCluster directly instead of guessing at whichever one happens to be Ready first.
Creating the CR can fail (a transient Kubernetes API error, for instance) without failing the whole request — the app row still exists in Postgres, a warning is logged, and the CR can be created later. This is a deliberate best-effort choice, not an oversight: a Postgres write is far cheaper to retry than a full GitOps round trip, and the app shows up in the API immediately either way.
None of this on its own means the app is actually running and reachable. KubeOperaApp.status.phase == "Ready" only means the manifest was generated and Flux picked it up — it says nothing about whether the container image pulled successfully, the pod passed its readiness probe, or the TLS certificate was actually issued. GET /api/v1/apps/{appId}/deploy-status is the endpoint that answers that real question: it composes the GitOps sync phase, live Pod/Deployment readiness (read from inside the tenant's vCluster, not the host cluster — see below), and, once both look healthy, an actual outbound HTTPS request against the computed subdomain. Only that last check proves DNS, the certificate, the Ingress, and the app itself are all genuinely working together; the wizard's own "Creating → Syncing → Starting → Verifying reachability" progress UI polls exactly this endpoint.
Every tenant operation targets the tenant's own vCluster, not the host cluster
Each tenant runs inside their own vCluster, provisioned by tenant-controller. Every per-app Kubernetes operation this API exposes — reading pods, streaming logs, listing events, scaling, restarting, computing deployment history — has to resolve which vCluster to talk to and build a client against it, rather than assuming a single shared cluster. That resolution logic lives in a small internal component (TenantClientResolver) that reads the tenant's exported kubeconfig Secret and caches the resulting client briefly to avoid re-reading it on every call. The same resolver backs the six vCluster-wide resource endpoints under /cloudspaces/{cloudSpaceId} (pods, deployments, services, ingresses, PVCs, jobs, cronjobs) that power the tenant dashboard's Workloads views, and the vCluster detail page.
Deployment history and rollback are read from Kubernetes, not a database
Every meaningful change to an app's pod template — a new image, a different port, an updated resource request — makes Kubernetes create a new ReplicaSet under the Deployment, each one carrying a deployment.kubernetes.io/revision annotation. GET /api/v1/apps/{appId}/deployments reads that ReplicaSet history directly rather than maintaining a parallel record in Postgres, which means it can never drift out of sync with what actually happened. A rollback (POST /api/v1/apps/{appId}/rollback) is, correspondingly, not a special operation — it reads the target ReplicaSet's image/port/env/resources, applies those values onto the app's current configuration, and pushes that through the exact same CR-update path a normal edit would use. Flux then reconciles the change and Kubernetes creates a new ReplicaSet reflecting the reverted state — there is no separate "undo" mechanism to keep correct.
REST API
Applications
| Method | Path | Description |
|---|---|---|
GET | /api/v1/apps | List applications |
POST | /api/v1/apps | Create an application (atomic — see above) |
GET | /api/v1/apps/{appId} | Application detail + live deployment status |
PUT | /api/v1/apps/{appId} | Update an application |
PATCH | /api/v1/apps/{appId}/domain | Change the app's domain |
DELETE | /api/v1/apps/{appId} | Delete an application (finalizer-cascades the CR, Flux Kustomization, and the running resources they applied) |
POST | /api/v1/apps/{appId}/restart | /scale | /stop | /start | Lifecycle actions, applied directly to the tenant vCluster (not GitOps-mediated — see note below) |
GET | /api/v1/apps/{appId}/metrics | /pods | /events | /logs | Live operational data, read from the tenant's own vCluster |
GET | /api/v1/apps/{appId}/deploy-status | The real "is it actually up" check described above |
GET | PUT | /api/v1/apps/{appId}/config | Full app configuration |
GET | PUT | /api/v1/apps/{appId}/config/env-vars | Environment variables only |
GET | /api/v1/apps/{appId}/deployments | Deployment (ReplicaSet) history |
GET | /api/v1/apps/{appId}/deployments/{deploymentId} | One revision's detail |
GET | /api/v1/apps/{appId}/deployments/{deploymentId}/export | Export a revision's config |
POST | /api/v1/apps/{appId}/deployments/{deploymentId}/tags | Tag a revision |
PUT | /api/v1/apps/{appId}/deployments/{deploymentId}/notes | Annotate a revision |
GET | /api/v1/apps/{appId}/deployments/compare | Compare two revisions |
POST | /api/v1/apps/{appId}/rollback | Roll back to a prior revision (see above) |
POST | /api/v1/apps/{appId}/clone | /promote | Clone or promote an app |
GET | PUT | /api/v1/apps/{appId}/deployment-strategy | Rolling/canary/blue-green strategy config |
GET | POST | /api/v1/apps/{appId}/canary (/promote, /abort, /pause, /resume) | Canary rollout control |
GET | POST | /api/v1/apps/{appId}/bluegreen (/switch, /rollback) | Blue/green rollout control |
Scale, restart, stop, and start are deliberately applied directly against the tenant clientset rather than round-tripped through a CR update and Flux's reconcile interval — users expect these to feel instant, and a multi-minute-default Flux interval would feel broken for them. The tradeoff is that runtime state from these actions can drift from the git-committed manifest until the next explicit configuration change syncs it back; this is an intentional choice, not an oversight.
Clusters, CloudSpaces & vClusters
| Method | Path | Description |
|---|---|---|
GET | /api/v1/clusters | List registered clusters |
GET | /api/v1/clusters/{clusterId} | Cluster detail |
POST | /api/v1/cloudspaces/{cloudSpaceId}/vclusters | Provision a vCluster |
GET | /api/v1/cloudspaces/{cloudSpaceId}/vclusters | List a CloudSpace's vClusters |
GET | /api/v1/cloudspaces/{cloudSpaceId}/vclusters/{name} | vCluster detail |
DELETE | /api/v1/cloudspaces/{cloudSpaceId}/vclusters/{name} | Delete a vCluster |
GET | /api/v1/cloudspaces/{cloudSpaceId}/vclusters/{name}/kubeconfig | Export kubeconfig |
GET | /api/v1/cloudspaces/{cloudSpaceId}/pods | /events | /deployments | /services | /ingresses | /pvcs | /jobs | /cronjobs | vCluster-wide resource lists (tenant dashboard Workloads view) |
GET | /api/v1/cloudspaces/{cloudSpaceId}/logs/{podName} | Pod logs, vCluster-wide |
CI/CD Pipelines & Kaniko builds
kubeopera-api is a trusted, tenant-aware gateway in front of two other services — cicd-gateway (pipelines) and build-service (Kaniko builds) — neither of which trusts a client-supplied tenant ID directly. Every route below resolves the caller's tenant_id from their JWT and forwards it explicitly on the internal, API-key-authenticated call, so a tenant can never see or act on another tenant's pipelines or builds by manipulating a request.
| Method | Path | Description |
|---|---|---|
GET | POST | /api/v1/pipelines | List / register a CI pipeline |
GET | PATCH | DELETE | /api/v1/pipelines/{id} | Pipeline detail, update, delete |
GET | /api/v1/pipelines/{id}/runs | Run history |
GET | /api/v1/pipelines/{id}/runs/summary | /api/v1/pipelines/summary | Per-pipeline / tenant-wide 7-day summary |
GET | POST | /api/v1/builds | List / start a Kaniko build |
GET | /api/v1/builds/{id} | /logs | Build status / logs |
GET | POST | /api/v1/build-configs | List / register an auto-rebuild-on-push config |
DELETE | /api/v1/build-configs/{id} | Remove a config |
POST | /api/v1/build-configs/{id}/trigger | Manually launch a build from a registered config, without waiting for a push |
Federation, onboarding & support
| Method | Path | Description |
|---|---|---|
GET | /api/v1/federation/overview | /health | /cost | /anomalies | /optimizer | Cross-cluster rollups for the platform-admin dashboard |
POST | /api/v1/onboarding/provision | Tenant self-service onboarding |
GET | /api/v1/onboarding/checklist | Onboarding checklist state |
GET | POST | /api/v1/support/tickets | Support ticket list / create |
GET | /api/v1/support/tickets/metrics | Support metrics |
PATCH | /api/v1/support/tickets/{id} | Update a ticket |
GET | /api/v1/incidents/active | Active incidents |
Environment Variables
| Variable | Description |
|---|---|
DATABASE_URL | PostgreSQL connection string |
AUTH_JWT_ACCESS_SECRET | Shared JWT signing secret (see auth-service) |
APP_CONTROLLER_ENABLED | true to enable KubeOperaApp CR creation — the GitOps path. Disabled only for local development. |
FLUX_GIT_URL / FLUX_GIT_BRANCH | Fleet (tenant manifest) repository and branch |
FLUX_GIT_SECRET_REF | Kubernetes Secret name holding Git credentials |
FLUX_TARGET_NAMESPACE | Default namespace for deployed applications |
FLUX_INTERVAL | Flux reconcile interval (default: 5m) |
CICD_GATEWAY_BASE_URL / CICD_GATEWAY_API_KEY | Internal, API-key-authenticated calls to cicd-gateway for the pipeline routes above |
BUILD_SERVICE_BASE_URL | Internal calls to build-service for the build routes above |