Skip to main content
Version: 2.0

Components

KubeOpera is a collection of focused components, each owning one job. This page is a guided tour: what each component does, what it talks to, and where to learn more. For how they fit together, start with Architecture.

GroupComponents
DashboardKubeOpera UI and its API gateway
Platform serviceskubeopera-api, auth-service, k8s-monitor, security, nodes-manager, cache
Automation servicesCI/CD gateway, anomaly detector, predictive scaler, incident manager
Application deliveryApp service, app-controller, build-service
Reactive AI pipelineObservability, analysis, action, feedback and recommendation agents
Reasoning runtimeagent-runtime and its agents
Observability servicesLogs, SLOs, APM, tracing, RCA, optimizer

Dashboard​

KubeOpera UI​

The web dashboard, built with the Next.js App Router, TypeScript and Tailwind CSS. The main areas are:

  • Dashboard — cluster health at a glance, with live metric cards.
  • Clusters — provision, connect and compare clusters.
  • Apps — create, deploy and observe applications.
  • Pipelines — CI/CD runs with stage-by-stage timelines.
  • Analytics — forecasts with confidence bands, anomalies and scaling approvals.
  • Incidents — incident list, detail and timeline.
  • AI Agents — pipeline health, and Agent Runs with a live view of each agent's reasoning and tool calls.
  • Monitoring — cost, performance, security, SLOs, APM and distributed traces.
  • Node Pools — Karpenter node pools and claims, with spot pricing.
  • AI Chat — a streaming assistant that can query every service.

Every backend call goes through the dashboard's server-side API routes (/api/*), which authenticate the caller, enforce the service allowlist and forward to the right microservice. See Architecture.

Platform services​

kubeopera-api​

The core application and cluster API. It manages cluster registration and lifecycle, application CRUD, deployments, pipelines and builds. Creating an app creates a KubeOperaApp custom resource for app-controller to deliver. → kubeopera-api

auth-service​

Identity for the whole platform: sign-in with OAuth2 providers (Google, GitHub, Keycloak) or email and password, JWT issuance, sessions and role-based access control. → auth-service

k8s-monitor​

Polls the Kubernetes API and cloud provider APIs on a regular cycle and produces:

  • a cluster health score (0–100) from node readiness, pod health and control-plane status;
  • cost per namespace and workload for AWS, GCP and on-premises;
  • right-sizing recommendations for over-provisioned and idle workloads;
  • node and pod metrics.

→ k8s-monitor

Security services​

A three-part scanning pipeline:

  • security-cron schedules scans;
  • security-collector scans cluster resources for vulnerabilities and misconfigurations and publishes findings to security.posture;
  • security-api stores findings and serves posture summaries, findings and compliance scores.

→ Security

nodes-manager​

Node lifecycle management built on Karpenter: node pools and claims, AI-assisted workload placement, spot market analysis, and proactive consolidation. It keeps node state in sync and publishes scaling events to nodes.events. → nodes-manager

Cache service​

A Redis-backed cache that keeps frequently read data fast and takes load off PostgreSQL. → Cache service

Automation services​

CI/CD gateway​

Receives GitHub and GitLab webhooks (verified with HMAC-SHA256), records pipelines, runs and stages, and publishes pipeline events to cicd.events for other services to react to. → CI/CD gateway

Anomaly detector​

Consumes anomaly events from k8s.anomalies, evaluates per-cluster alert rules, runs self-healing actions that are approved to run automatically, logs every remediation, and publishes the result to k8s.selfheal. → Anomaly detector

Predictive scaler​

Forecasts load from historical metrics using Holt-Winters smoothing, with 95% confidence bands, and proposes scaling decisions that you approve from the Analytics page or through an AI agent. It needs at least 24 data points before its first forecast. → Predictive scaler

Incident manager​

Manages incidents from detection to resolution:

  • consumes anomalies, self-healing results and security findings;
  • groups related events (same cluster, category and namespace within five minutes) into one incident to prevent alert storms;
  • runs runbooks (kubectl, HTTP, notify and wait steps);
  • notifies Slack and PagerDuty;
  • tracks MTTR, MTBF and SLA compliance.

→ Incident manager

Application delivery​

App service​

The manifest generator. Given an app definition — name, image, ports, resources, autoscaling, ingress — it produces production-ready Kubernetes manifests. → App lifecycle

app-controller​

The operator that reconciles KubeOperaApp resources: it resolves the tenant's vCluster, generates manifests through the app service, commits them to the fleet Git repository, and makes sure Flux inside the vCluster picks them up. Deleting the app cleans everything up. → app-controller

build-service​

Builds container images from a Git repository with Kaniko, so apps can be deployed straight from source. → build-service

Reactive AI pipeline​

A chain of small services connected by RabbitMQ that detects, acts and learns. → Reactive AI Pipeline

AgentRole
Observability agentStores telemetry snapshots and publishes them to observability.telemetry.
Analysis agentRuns Z-score detection over a 60-point rolling window with adaptive per-cluster thresholds; publishes decisions and insights; tunes thresholds from feedback.signals.
Action agentExecutes auto-approved decisions — scale a deployment, restart a pod, cordon a node — and publishes each outcome.
Feedback agentCompares post-action telemetry with the pre-action baseline and publishes a reinforcement or correction signal.
Recommendation agentTurns insights into structured, prioritized recommendations using Claude.

Reasoning runtime​

agent-runtime runs Claude-powered agents that reason step by step and call live tools (the ReAct pattern). Each agent is specialized:

AgentFocus
SRE OrchestratorEnd-to-end investigation with access to every tool.
Security AuditorPosture, findings and RBAC.
Cost OptimizerSpend, waste and right-sizing.
Incident ResponderTriage and remediation of active incidents.
App AdvisorPer-application best practices and improvements.
Load Test AnalystPerformance and load-test results.
NodeOpsNode pools, capacity and consolidation.

Every tool call is streamed to the UI and stored, so each run has a complete, reviewable history. The SRE Orchestrator launches automatically when the analysis risk score exceeds 70. → Agent runtime

Observability services​

ServiceWhat it does
log-gatewayTranslates service and namespace queries into LogQL and returns log lines and extracted patterns from Loki.
slo-managerStores SLO definitions, evaluates them against Prometheus, and reports error budget and burn rate.
apm-gatewayServes per-service request rate, p99 latency and error rate from Prometheus.
tracing-gatewayServes trace detail, service dependency maps and slow/error trace search from Jaeger.
rca-engineGathers anomalies, logs and incidents in parallel and asks Claude for a structured root cause with contributing factors and confidence.
Optimizer (kubeopera-ai)Continuously assesses cluster health, offers on-demand optimization and load-test analysis, and publishes degradation signals to k8s.anomalies for the anomaly detector.

→ Observability services