Skip to main content
Version: 1.0

Technical Overview

Purpose​

KubeOpera is built on Kubernetes' own core primitives and leverages a number of established industry tools and solutions to provide a complete solution for managing, monitoring, troubleshooting, optimising, and deploying applications to Kubernetes environments. It eliminates the need for multiple disparate tools, reduces complexity, and lowers the operational burden on DevOps and platform engineering teams.

Core Principles​

Event-driven​

Services communicate through RabbitMQ topic exchanges. When the observability agent collects telemetry, the analysis agent receives it immediately and propagates decisions downstream — no polling, no manual coordination.

Autonomous with human override​

Pod restarts, node cordons, and replica scaling happen automatically based on analysis decisions. Every action is logged and feedback is collected. Operators can require manual approval on any rule via AutoApprove: false.

AI as a first-class operator​

The agent-runtime runs Claude-powered reasoning agents that use over 50 (and counting) specialized live tools to investigate clusters, correlate signals, and produce prioritised recommendations. Agents can be triggered manually or auto-invoked when analysis risk scores exceed 70/100.

Fully observable AI​

Every tool call an AI agent makes is streamed as a Server-Sent Event to the UI and persisted to PostgreSQL. You can watch the agent reason in real time and review the complete tool call history after the fact.

Key Capabilities​

AreaWhat it does
Multi-cluster managementProvision, register, and operate clusters across AWS and GCP from one UI
Cost intelligenceReal-time per-namespace/workload cost tracking plus Holt-Winters forecasts
Security postureContinuous scanning pipeline with compliance scoring and RBAC analysis
CI/CD visibilityGitHub/GitLab webhook tracking with stage-level pipeline timelines
Anomaly detectionZ-score analysis over rolling windows with configurable thresholds
Self-healingAutomated pod restarts, node cordons, and scale-outs with action logging
Predictive scalingHPA recommendations from load forecasts with one-click approval
Incident managementFull lifecycle with runbooks, correlation dedup, Slack/PagerDuty notifications
App lifecycleSchema-driven manifest generation (kubegenic) + Flux GitOps deployment via CRD
Agentic AIFour LLM reasoning agents with tool use, streaming, and interleaved thinking
MCP integration49 tools across five toolsets for Claude Desktop, Claude Code, and other MCP clients

Technology Stack​

LayerTechnologyPurpose
FrontendNext.js, TypeScript, TailwindApp Router, SSE streaming, server components
BackendGo (hexagonal architecture)Performance, single-binary deployment
Message busRabbitMQ topic exchangesFan-out without coupling
AI SDKanthropic-sdk-go v1.37BetaToolRunnerStreaming, interleaved thinking
ForecastingHolt-Winters (pure Go)No Python or ML framework dependencies
CI/CDFlux v2 CRDs via controller-runtimeDeclarative, Kubernetes-native reconciliation
MCPFork of kubernetes-mcp-serverStdio by default, HTTP/SSE also supported
OrchestrationvClusterPer-tenant isolation within a shared cluster