Components
KubeOpera is a collection of focused components, each owning one job. This page is a guided tour: what each component does, what it talks to, and where to learn more. For how they fit together, start with Architecture.
| Group | Components |
|---|---|
| Dashboard | KubeOpera UI and its API gateway |
| Platform services | kubeopera-api, auth-service, k8s-monitor, security, nodes-manager, cache |
| Automation services | CI/CD gateway, anomaly detector, predictive scaler, incident manager |
| Application delivery | App service, app-controller, build-service |
| Reactive AI pipeline | Observability, analysis, action, feedback and recommendation agents |
| Reasoning runtime | agent-runtime and its agents |
| Observability services | Logs, SLOs, APM, tracing, RCA, optimizer |
Dashboard
KubeOpera UI
The web dashboard, built with the Next.js App Router, TypeScript and Tailwind CSS. The main areas are:
- Dashboard — cluster health at a glance, with live metric cards.
- Clusters — provision, connect and compare clusters.
- Apps — create, deploy and observe applications.
- Pipelines — CI/CD runs with stage-by-stage timelines.
- Analytics — forecasts with confidence bands, anomalies and scaling approvals.
- Incidents — incident list, detail and timeline.
- AI Agents — pipeline health, and Agent Runs with a live view of each agent's reasoning and tool calls.
- Monitoring — cost, performance, security, SLOs, APM and distributed traces.
- Node Pools — Karpenter node pools and claims, with spot pricing.
- AI Chat — a streaming assistant that can query every service.
Every backend call goes through the dashboard's server-side API routes (/api/*), which authenticate the caller, enforce the service allowlist and forward to the right microservice. See Architecture.
Platform services
kubeopera-api
The core application and cluster API. It manages cluster registration and lifecycle, application CRUD, deployments, pipelines and builds. Creating an app creates a KubeOperaApp custom resource for app-controller to deliver. → kubeopera-api
auth-service
Identity for the whole platform: sign-in with OAuth2 providers (Google, GitHub, Keycloak) or email and password, JWT issuance, sessions and role-based access control. → auth-service
k8s-monitor
Polls the Kubernetes API and cloud provider APIs on a regular cycle and produces:
- a cluster health score (0–100) from node readiness, pod health and control-plane status;
- cost per namespace and workload for AWS, GCP and on-premises;
- right-sizing recommendations for over-provisioned and idle workloads;
- node and pod metrics.
Security services
A three-part scanning pipeline:
- security-cron schedules scans;
- security-collector scans cluster resources for vulnerabilities and misconfigurations and publishes findings to
security.posture; - security-api stores findings and serves posture summaries, findings and compliance scores.
→ Security
nodes-manager
Node lifecycle management built on Karpenter: node pools and claims, AI-assisted workload placement, spot market analysis, and proactive consolidation. It keeps node state in sync and publishes scaling events to nodes.events. → nodes-manager
Cache service
A Redis-backed cache that keeps frequently read data fast and takes load off PostgreSQL. → Cache service
Automation services
CI/CD gateway
Receives GitHub and GitLab webhooks (verified with HMAC-SHA256), records pipelines, runs and stages, and publishes pipeline events to cicd.events for other services to react to. → CI/CD gateway
Anomaly detector
Consumes anomaly events from k8s.anomalies, evaluates per-cluster alert rules, runs self-healing actions that are approved to run automatically, logs every remediation, and publishes the result to k8s.selfheal. → Anomaly detector
Predictive scaler
Forecasts load from historical metrics using Holt-Winters smoothing, with 95% confidence bands, and proposes scaling decisions that you approve from the Analytics page or through an AI agent. It needs at least 24 data points before its first forecast. → Predictive scaler
Incident manager
Manages incidents from detection to resolution:
- consumes anomalies, self-healing results and security findings;
- groups related events (same cluster, category and namespace within five minutes) into one incident to prevent alert storms;
- runs runbooks (kubectl, HTTP, notify and wait steps);
- notifies Slack and PagerDuty;
- tracks MTTR, MTBF and SLA compliance.
Application delivery
App service
The manifest generator. Given an app definition — name, image, ports, resources, autoscaling, ingress — it produces production-ready Kubernetes manifests. → App lifecycle
app-controller
The operator that reconciles KubeOperaApp resources: it resolves the tenant's vCluster, generates manifests through the app service, commits them to the fleet Git repository, and makes sure Flux inside the vCluster picks them up. Deleting the app cleans everything up. → app-controller
build-service
Builds container images from a Git repository with Kaniko, so apps can be deployed straight from source. → build-service
Reactive AI pipeline
A chain of small services connected by RabbitMQ that detects, acts and learns. → Reactive AI Pipeline
| Agent | Role |
|---|---|
| Observability agent | Stores telemetry snapshots and publishes them to observability.telemetry. |
| Analysis agent | Runs Z-score detection over a 60-point rolling window with adaptive per-cluster thresholds; publishes decisions and insights; tunes thresholds from feedback.signals. |
| Action agent | Executes auto-approved decisions — scale a deployment, restart a pod, cordon a node — and publishes each outcome. |
| Feedback agent | Compares post-action telemetry with the pre-action baseline and publishes a reinforcement or correction signal. |
| Recommendation agent | Turns insights into structured, prioritized recommendations using Claude. |
Reasoning runtime
agent-runtime runs Claude-powered agents that reason step by step and call live tools (the ReAct pattern). Each agent is specialized:
| Agent | Focus |
|---|---|
| SRE Orchestrator | End-to-end investigation with access to every tool. |
| Security Auditor | Posture, findings and RBAC. |
| Cost Optimizer | Spend, waste and right-sizing. |
| Incident Responder | Triage and remediation of active incidents. |
| App Advisor | Per-application best practices and improvements. |
| Load Test Analyst | Performance and load-test results. |
| NodeOps | Node pools, capacity and consolidation. |
Every tool call is streamed to the UI and stored, so each run has a complete, reviewable history. The SRE Orchestrator launches automatically when the analysis risk score exceeds 70. → Agent runtime
Observability services
| Service | What it does |
|---|---|
| log-gateway | Translates service and namespace queries into LogQL and returns log lines and extracted patterns from Loki. |
| slo-manager | Stores SLO definitions, evaluates them against Prometheus, and reports error budget and burn rate. |
| apm-gateway | Serves per-service request rate, p99 latency and error rate from Prometheus. |
| tracing-gateway | Serves trace detail, service dependency maps and slow/error trace search from Jaeger. |
| rca-engine | Gathers anomalies, logs and incidents in parallel and asks Claude for a structured root cause with contributing factors and confidence. |
| Optimizer (kubeopera-ai) | Continuously assesses cluster health, offers on-demand optimization and load-test analysis, and publishes degradation signals to k8s.anomalies for the anomaly detector. |