Core Services
KubeOpera's backend is a set of focused Go services. Each owns one domain, its own database schema and its own API, and they cooperate through REST calls and RabbitMQ events. This page explains what they have in common, lists them all, and maps how events flow between them.
Service catalog
| Group | Service | Port | What it does |
|---|---|---|---|
| Platform | kubeopera-api | 8090 | Apps, deployments, builds, CloudSpaces, clusters, federation. |
| auth-service | 8081 | Identity, tokens, SSO, RBAC and credential storage. | |
| k8s-monitor | 8085 | Health scores, cost, metrics and anomaly signals. | |
| Security | 8086 | Vulnerability scanning, posture and compliance. | |
| cicd-gateway | 8087 | CI/CD webhooks, pipelines and runs. | |
| nodes-manager | 8115 | Node pools, Karpenter, spot and consolidation. | |
| cache-service | 8080 | Shared caching and rate-limit counters. | |
| Automation | anomaly-detector | 8088 | Alert rules and self-healing. |
| predictive-scaler | 8089 | Load forecasts and scaling decisions. | |
| incident-manager | 8090 | Incidents, correlation, runbooks and notifications. | |
| App delivery | App lifecycle | 8112 | Manifest generation from an app schema. |
| app-controller | — | Reconciles KubeOperaApp resources through GitOps. | |
| build-service | 8098 | Container image builds from source. | |
| app-advisor-srv | 8105 | Per-app profiles, advice and live metrics. | |
| advisor-controller | — | Runs an App Advisor in each tenant's vCluster. | |
| Observability | log-gateway | 8099 | Logs and log patterns (Loki). |
| slo-manager | 8100 | SLOs, error budgets and burn rates. | |
| apm-gateway | 8101 | Application performance metrics (Prometheus). | |
| tracing-gateway | 8102 | Traces and service maps (Jaeger). | |
| rca-engine | 8103 | AI root-cause analysis. | |
| Optimization | kubeopera-ai (Optimizer) | 8113 | Continuous AI health and optimization analysis. |
| k8s-optimizer | 8097 | ML-based resource right-sizing. | |
| Infrastructure | tenant-controller | — | Reconciles Tenant resources into CloudSpaces. |
| host-cluster-controller | — | Reconciles HostCluster resources. | |
| cluster-provisioner | — | Provisions cloud clusters with Terraform. | |
| gitops-scaffolder | — | Generates GitOps environments for new clusters. |
The reactive AI pipeline and agent runtime are documented in AI Agents: observability-agent-srv (8092), analysis-agent-srv (8093), action-agent-srv (8094), feedback-agent-srv (8095), recommendation-agent-srv (8096) and agent-runtime (8111).
Service port reference
Use these ports with kubectl port-forward when you need to reach a service directly:
kubectl port-forward -n kubeopera-core svc/k8s-monitor 8085:8085
curl localhost:8085/healthz
Two services share port 8090 (kubeopera-api and incident-manager); they have different Service names, so forward each to a different local port if you need both.
How every service is built
Services follow hexagonal architecture — business logic in the middle, with adapters for HTTP, messaging and storage around it — so every service is laid out the same way:
cmd/server/main.go Entry point: wire dependencies, start the server
internal/
core/
domain/ Domain types
ports/ Interfaces: repositories, publishers, clients
services/ Business logic (depends only on ports)
adapters/
http/ Router, handlers, auth middleware, DTOs
messaging/rabbitmq/ Consumers and publishers
repository/postgres/ Database repositories
migrations/ SQL migrations for the service's own schema
deployments/
base/ Kubernetes manifests
overlays/{development,staging,production}/
Dockerfile
Once you've read one service, you can find your way around any of them.
Shared conventions
| Concern | Convention |
|---|---|
| HTTP | chi router; JSON APIs under /api/v1; GET /healthz for liveness and readiness; GET /metrics for Prometheus; OpenAPI at /openapi.json. |
| Auth | Every API requires authentication through the shared auth middleware: user or service tokens issued by auth-service. |
| Data | Each service owns a PostgreSQL schema; no foreign keys or queries across schemas. Migrations run on startup and must succeed before the service reports ready. |
| Messaging | Durable RabbitMQ topic exchanges; queue names {consumer}.{exchange}; failed messages retried, then dead-lettered to {queue}.dlq. |
| Configuration | Environment variables; required variables are validated at startup. |
| Observability | Structured JSON logs, Prometheus metrics, OpenTelemetry traces. |
Events
Services publish what they learn to RabbitMQ topic exchanges, and any service can subscribe. This is how KubeOpera grows: new capabilities subscribe to existing events without changing the services that publish them.
| Exchange | Routing keys | Publisher | Consumers |
|---|---|---|---|
observability.telemetry | telemetry.{cluster} | observability-agent-srv | analysis-agent-srv |
analysis.decisions | decision.{type} | analysis-agent-srv | action-agent-srv |
analysis.insights | insight.{severity} | analysis-agent-srv | recommendation-agent-srv, agent-runtime |
action.outcomes | outcome.{status} | action-agent-srv | feedback-agent-srv |
feedback.signals | signal.{type} | feedback-agent-srv | analysis-agent-srv |
k8s.anomalies | anomaly.{source}.{severity} | k8s-monitor, kubeopera-ai | anomaly-detector, incident-manager |
k8s.selfheal | selfheal.{status} | anomaly-detector | incident-manager |
security.posture | posture.{severity} | security-collector | incident-manager |
cicd.events | cicd.{provider}.{event} | cicd-gateway | observability-agent-srv, incident-manager |
incidents.events | incident.{event} | incident-manager | kubeopera-api (webhooks), agent-runtime |
nodes.events | nodes.{event} | nodes-manager | incident-manager, observability-agent-srv |
Routing keys start with a fixed segment per exchange, and consumers bind with a pattern such as anomaly.#, so every publisher's messages reach every consumer of that exchange.
Next steps
- Configuration — how services are configured.
- Monitoring — watching the services themselves.
- Architecture — how the services fit into the platform.