Troubleshooting Guide
Problems are grouped by what you're seeing. For each one: what to check, and how to fix it.
Most problems can be narrowed down in a minute with:
flux get kustomizations # is Git applied?
kubectl get pods -n kubeopera-core # is everything running?
kubectl logs deploy/<service> -n kubeopera-core --tail=100
Or ask the AI Chat — "what's wrong with the analysis agent?" — which checks health, logs and recent events for you.
Installation and upgrades
A Flux Kustomization stays not-ready
Check: flux get kustomizations — the message says why.
- "dependency not ready" — something it depends on hasn't finished. Fix the first failing Kustomization in the chain.
- "failed to decrypt" on a sealed secret — the secret was sealed for a different cluster. Re-seal it with this cluster's key (
kubeseal --fetch-cert). - Certificate or ingress errors — check the next section.
The dashboard address doesn't load, or has a certificate error
Check:
kubectl get certificate -A # READY should be True
kubectl describe certificate <name> -n kubeopera-core
kubectl get ingress -n kubeopera-core
- A certificate stuck not-ready usually means the DNS-01 challenge can't complete: check that external-dns and cert-manager have permission to update your DNS zone.
- If DNS for your domain doesn't point at the ingress load balancer yet, use
kubectl port-forwardto reach the dashboard in the meantime.
A service keeps restarting after an upgrade
Check its logs for a startup error. Services fail fast on missing required configuration or a failed database migration, and say which. Fix the setting in the service's overlay, or see the migration error for the table or column involved.
Signing in
"Invalid credentials" for a user who is sure of their password
- They may be required to sign in with SSO — check Settings → Single Sign-On for their organization.
- Too many failed attempts temporarily lock sign-in; wait for the lock to expire or reset the password.
Sign-in succeeds but pages show "unauthorized"
The user's role doesn't include the permission the page needs. Check their roles in Settings → Users. For tenant users, confirm they belong to the right tenant.
Apps and deployments
An app is stuck in "Syncing"
Check the app's Events tab, then its GitOps status:
- A validation error in the generated manifests appears as a Flux error on the app — fix the configuration and redeploy.
- If the CloudSpace is not
Ready, fix that first (see below).
An app is stuck in "Starting"
The pods aren't becoming ready. Open the app's Pods and Logs tabs:
- ImagePullBackOff — the image name is wrong, or it's private and needs registry credentials.
- CrashLoopBackOff — the app is failing at startup; its logs say why.
- Pending — the CloudSpace's quota is exhausted or the cluster is out of capacity; the event message says which.
- Running but not ready — check the app's port matches what it listens on, and its readiness path.
An app is ready but "View live app" never appears
KubeOpera is waiting for the app to answer over HTTPS. On a first deploy, certificate issuance takes a few minutes. If it takes longer, check the app's certificate in its Events tab, and that the app responds on its configured port.
A manual kubectl edit keeps being reverted
That's GitOps doing its job — Git is the source of truth. Make the change in the dashboard (or in Git), and it will stick.
CloudSpaces
A CloudSpace is stuck in "Provisioning" or shows "Failed"
Check the reason on the CloudSpace page, or:
kubectl get tenant -n kubeopera-system
kubectl describe tenant <name> -n kubeopera-system
Provisioning retries automatically. Common causes are the host cluster running out of capacity, or a temporary failure pulling the vCluster chart. Once resolved, select Retry (or wait for the next automatic retry).
Dashboard
A page shows no data
- Check the cluster/CloudSpace selector in the top bar.
- For latency, network and error charts, your apps need to expose Prometheus metrics — see Monitoring your applications.
- If one service's data is missing everywhere, check that service is running.
A newly added integration can't reach its service
The dashboard only calls allowlisted service hostnames. If you run a service under a custom hostname, add it to ALLOWED_UPSTREAM_HOSTS (see Configuration).
The AI pipeline
No recommendations or agent runs are appearing
- Check an AI provider key is configured (Settings → AI Provider Key) and the tenant hasn't reached its quota — the Recommendations page says if generation is paused.
- Check the pipeline's status on
/agents— every agent should be healthy. - Recommendations only appear when something notable happens; a quiet cluster produces few.
The same automatic action keeps happening
KubeOpera learns from outcomes: if an action doesn't help, feedback raises the threshold and the action becomes less likely. If you want it to stop immediately, change the cluster's decision rules so the action requires approval.
Events are published but a consumer never receives them
RabbitMQ topic exchanges route by routing key. Check the Message bus dashboard in Grafana for unroutable messages, and compare the publisher's routing key with the consumer's binding pattern.
Nodes and clusters
A node is Ready but SchedulingDisabled
The node was cordoned — by a person, by an automated action, or by the node termination handler in response to a cloud spot-capacity warning. nodes-manager uncordons nodes automatically when the warning passes without an interruption. To check who cordoned it, look at the node's events and the action log on /agents/actions.
A cluster is stuck in a setup phase
Open the cluster's card for the current phase and message, or:
kubectl describe hostcluster <cluster-id> -n kubeopera-system
kubectl logs -n kubeopera-provisioning job/provision-<cluster-id> # for Create New
See Clusters: the phases for what each phase does.
Still stuck?
- Search GitHub Discussions or ask a question.
- Open an issue on GitHub with the output of the "start here" commands.
- Customers can open a ticket from Support in the dashboard.