Technical Overview
Purpose
KubeOpera is built on Kubernetes' own core primitives and leverages a number of established industry tools and solutions to provide a complete solution for managing, monitoring, troubleshooting, optimising, and deploying applications to Kubernetes environments. It eliminates the need for multiple disparate tools, reduces complexity, and lowers the operational burden on DevOps and platform engineering teams.
Core Principles
Event-driven
Services communicate through RabbitMQ topic exchanges. When the observability agent collects telemetry, the analysis agent receives it immediately and propagates decisions downstream — no polling, no manual coordination.
Autonomous with human override
Pod restarts, node cordons, and replica scaling happen automatically based on analysis decisions. Every action is logged and feedback is collected. Operators can require manual approval on any rule via AutoApprove: false.
AI as a first-class operator
The agent-runtime runs Claude-powered reasoning agents that use over 50 (and counting) specialized live tools to investigate clusters, correlate signals, and produce prioritised recommendations. Agents can be triggered manually or auto-invoked when analysis risk scores exceed 70/100.
Fully observable AI
Every tool call an AI agent makes is streamed as a Server-Sent Event to the UI and persisted to PostgreSQL. You can watch the agent reason in real time and review the complete tool call history after the fact.
Key Capabilities
| Area | What it does |
|---|---|
| Multi-cluster management | Provision, register, and operate clusters across AWS and GCP from one UI |
| Cost intelligence | Real-time per-namespace/workload cost tracking plus Holt-Winters forecasts |
| Security posture | Continuous scanning pipeline with compliance scoring and RBAC analysis |
| CI/CD visibility | GitHub/GitLab webhook tracking with stage-level pipeline timelines |
| Anomaly detection | Z-score analysis over rolling windows with configurable thresholds |
| Self-healing | Automated pod restarts, node cordons, and scale-outs with action logging |
| Predictive scaling | HPA recommendations from load forecasts with one-click approval |
| Incident management | Full lifecycle with runbooks, correlation dedup, Slack/PagerDuty notifications |
| App lifecycle | Schema-driven manifest generation (kubegenic) + Flux GitOps deployment via CRD |
| Agentic AI | Four LLM reasoning agents with tool use, streaming, and interleaved thinking |
| MCP integration | 49 tools across five toolsets for Claude Desktop, Claude Code, and other MCP clients |
Technology Stack
| Layer | Technology | Purpose |
|---|---|---|
| Frontend | Next.js, TypeScript, Tailwind | App Router, SSE streaming, server components |
| Backend | Go (hexagonal architecture) | Performance, single-binary deployment |
| Message bus | RabbitMQ topic exchanges | Fan-out without coupling |
| AI SDK | anthropic-sdk-go v1.37 | BetaToolRunnerStreaming, interleaved thinking |
| Forecasting | Holt-Winters (pure Go) | No Python or ML framework dependencies |
| CI/CD | Flux v2 CRDs via controller-runtime | Declarative, Kubernetes-native reconciliation |
| MCP | Fork of kubernetes-mcp-server | Stdio by default, HTTP/SSE also supported |
| Orchestration | vCluster | Per-tenant isolation within a shared cluster |