Troubleshooting Guide
The problems on this page are real, recurring failure classes specific to how KubeOpera is built — a proxy-heavy Next.js frontend, ~35 independently-deployed Go services, and a GitOps deployment model — rather than generic Kubernetes advice you could find anywhere. Each one below has actually happened and been root-caused; they're grouped by symptom, since that's usually what you're starting from.
A tenant-facing page 403s or shows "no resources found," but works for an admin
The frontend never lets the browser talk to a backend service directly — every call goes through a Next.js API route acting as a server-side proxy, and several of those routes need to resolve which tenant's Kubernetes cluster (technically, vCluster) to actually query. When that resolution logic calls back into the frontend's own other API routes, it has to use the app's internal loopback address (APP_INTERNAL_URL, defaulting to http://localhost:3000) rather than the request's own public hostname — a pod calling its own externally-facing URL depends on ingress hairpin routing, which isn't guaranteed to work and can fail 100% of the time, silently, with no error logged at either call site. If a route works fine for a super_admin session (who isn't subject to this tenant-resolution step at all) but consistently fails for real tenant/customer sessions specifically, this pattern — grep the route for nextUrl.origin — is the first thing worth checking, before assuming it's a permissions bug.
A backend request from the frontend fails with no useful error
Every backend hostname a Next.js API route is allowed to call has to be explicitly listed in an SSRF allowlist (app/api/_lib/upstream.ts). Adding a new backend integration without adding its hostname there doesn't fail loudly — the fetch is simply blocked. If a newly-added proxy route can't reach its backend at all, checking this allowlist before debugging the backend service itself will usually save time.
A service's database migrations silently never ran
Two independent ways this has happened, both worth checking: the multi-stage Docker build copying the compiled binary into the final image but forgetting to also copy the migrations/ directory it needs at runtime (the migration tool finds nothing to run, logs a warning, and the service starts anyway against an unmigrated database) — and, in a shared Postgres instance where several services' schemas coexist, the migration tool's own bookkeeping table (goose_db_version) resolving against a stray, differently-schemaed copy of itself via Postgres's search_path fallthrough, silently skipping every real migration. If a service's logs show a relation "..." does not exist error for a table you're confident the code creates, check that its image actually contains its migration files, and that its migration bookkeeping table is explicitly schema-qualified.
A gitops deploy rolled out the wrong image on a container that isn't the main one
Several services run a database migration as an init container, separate from the main application container — and the automated pipeline that bumps a service's deployed image tag on every release only ever updates the main container's image reference, never the init container's. Left unnoticed, this means a service can run its new application code against migrations still pinned to an old release, sometimes for several releases in a row before anyone notices. Whenever you're deploying a service that has a migrate init container, it's worth explicitly confirming both image references were bumped together, not just the main one.
Two services publish to the same RabbitMQ exchange, but one's messages never arrive
All of KubeOpera's message exchanges are topic type, which means a message's routing key has to match a consumer's binding pattern segment-by-segment from the left — agreeing on the exchange name alone isn't enough. A publisher using a routing key like service-name.severity and a consumer bound to a pattern like something-else.# will never exchange a single message, and nothing about this fails loudly on either side; the broker just drops what it can't route. If a service's messages seem to vanish with no consumer-side error at all, comparing the exact publish routing key against the exact binding pattern, on both ends, is the fastest way to find this class of bug — several real examples of it are called out on the Backend Services overview page.
A JSON/JSONB column write fails with "invalid input syntax for type json"
This shows up on services using the bun ORM: a Go []byte field holding pre-marshaled JSON needs an explicit type:jsonb tag alongside its column name, or the ORM defaults to treating it as a plain binary (bytea) column — which Postgres then rejects against a real jsonb column the migration actually created. If this error appears on a write that looks like it should obviously work, checking the struct tag for type:jsonb before assuming the JSON payload itself is malformed will usually be faster.
A node is Ready but stuck SchedulingDisabled for no apparent reason
aws-node-termination-handler cordons a node in response to an AWS EC2 Spot rebalance recommendation — a real but only elevated-risk signal, not a guaranteed interruption — and, by this cluster's own configuration, never reverses that cordon on its own if no actual interruption follows. nodes-manager's auto-uncordon reconciler exists specifically to catch and reverse this pattern automatically after a grace period; if a node is stuck like this for longer than that grace period, checking whether the reconciler is actually running (and whether the node's only taint really is the bare unschedulable one, versus something added manually or by Karpenter) is the right next step before uncordoning it by hand.