Skip to main content
Version: 1.0

Optimizer

Optimizer (service name kubeopera-ai) is the AI operations agent. It continuously monitors cluster health, provides on-demand resource optimisation and load test analysis powered by Claude, and offers an interactive kubectl/Prometheus troubleshooting conversation. It is one of five services (alongside agent-runtime, recommendation-agent-srv, app-advisor-srv, and the frontend's own AI chat) that resolve their Anthropic API key per request rather than from a static environment variable — see Auth Service for how that resolution works. There is no platform-wide ANTHROPIC_API_KEY anymore; a tenant's own key, if configured, is always used ahead of the platform's shared key.

Key Capabilities​

Every monitoring cycle (INTERVAL, default 5 minutes), the service gathers node status, pod state, recent events, and resource usage, and asks Claude for a structured health assessment — a 0–100 health_score, an overall status, and a list of severity-ranked issues, each with a concrete fix command where one applies. That same analysis is available on demand through POST /api/v1/analyze/health, independent of the background cycle.

Resource optimisation works similarly but on a longer horizon: it compares seven days of average CPU and memory usage (pulled from apm-gateway, never via kubectl port-forward) against each workload's current requests and limits, and produces a Markdown report with ready-to-run kubectl patch commands and an estimated cost saving. Load test analysis takes a test's output and resource snapshots and returns a p95/p99 assessment, likely bottlenecks, scaling recommendations, and a prioritised action list.

The interactive troubleshooter (POST /api/v1/troubleshoot) runs a bounded, multi-turn conversation — capped at ten iterations, with history pruned to the last twenty turn-pairs so the context window doesn't grow unbounded — in which Claude can call two intentionally narrow tools: execute_kubectl, which runs via exec.Command with arguments passed as a Go string slice (never through a shell, so there's no injection surface), and query_prometheus, a plain HTTP GET against the configured APM gateway.

Continuous Monitoring & RabbitMQ Integration​

When RABBITMQ_URL is set, each cycle publishes critical and warning issues to the k8s.anomalies topic exchange, with routing key kubeopera-ai.{severity}.

This publish currently reaches no consumer. anomaly-detector, the service that would otherwise pick these events up, binds its queue to the pattern anomaly.# — and on an AMQP topic exchange, the first routing-key segment has to match literally. kubeopera-ai.critical never matches anomaly.#, so every one of these messages is silently dropped by the broker; nothing is malfunctioning loudly, there's just no delivery. The rest of the intended chain — anomaly-detector to incident-manager to an agent-runtime auto-triggered SRE Orchestrator investigation once cumulative risk crosses 70 — is real, working code on the receiving end, but never fires today because nothing reaches it from this service. Fixing this is a routing-key change on one side or an additional binding on the other; it has not shipped yet.

The published AnomalyMessage payload, for when this is fixed:

{
"cluster_id": "kubeopera-prod",
"metric": "payments-api",
"value": 32.0,
"z_score": 3.0,
"baseline": 68.0,
"severity": "critical",
"message": "Multiple pods in CrashLoopBackOff, memory pressure detected"
}

Domain Model​

HealthAnalysis
├── health_score: 0-100
├── status: healthy | warning | critical
├── summary
├── issues: []HealthIssue
│ ├── severity: critical | warning | info
│ ├── component, description, recommendation
│ └── command (kubectl fix command, if applicable)
└── generated_at

OptimizationResult
├── recommendations (Markdown report with kubectl patch commands)
└── generated_at

LoadTestResult
├── analysis (Markdown report with numbered action list)
├── results_path
└── generated_at

StatusResponse ← GET /api/v1/status
├── last_health
├── last_optimize
├── namespace
└── uptime

REST API​

MethodPathDescription
GET/healthzLiveness probe
GET/api/v1/statusLast health result, last optimize result, uptime
POST/api/v1/analyze/healthOn-demand health analysis
POST/api/v1/analyze/optimizeOn-demand resource optimisation
POST/api/v1/analyze/load-testLoad test analysis ({ "results_path": "..." }, optional)
POST/api/v1/troubleshootMulti-turn kubectl/Prometheus Q&A ({ "query": "..." })

Integration with agent-runtime​

Two tools are registered in agent-runtime and exposed to all agents: get_optimization_status (fetches GET /api/v1/status) and get_load_test_analysis (triggers POST /api/v1/analyze/load-test). A dedicated Load Test Analyst agent (load_test_analyst) combines these with get_cluster_health, get_pod_metrics, get_node_metrics, get_scaling_forecasts, and get_anomaly_events to produce a full performance report with a prioritised remediation list.

These two tools are the only real frontend/agent-facing surface this service currently has — the health-analysis issue list, the resource-optimisation report, and the interactive troubleshooter described above are otherwise reachable only by calling the API directly; no dedicated dashboard page exists for them yet.

Environment Variables​

VariableDefaultDescription
AUTH_JWT_ACCESS_SECRET—Required. Shared JWT signing secret — the process exits at startup without it.
NAMESPACEkubeopera-prodKubernetes namespace to monitor
PORT8113HTTP port (the deployed default; some historical documentation and the code's own fallback still say 8104 — treat 8113 as the live value)
INTERVAL5mMonitoring cycle (e.g. 5m, 10m, 1h)
APM_GATEWAY_BASE_URLhttp://localhost:8101Prometheus/APM-gateway for 7-day metric queries
RABBITMQ_URL—Optional. Enables anomaly event publishing to k8s.anomalies — see the caveat above before relying on it reaching a consumer
KUBECONFIGin-clusterPath to kubeconfig for local development
AUTH_SERVICE_BASE_URL / AI_CREDENTIAL_INTERNAL_API_KEY—Used to resolve the Anthropic key per request — see Auth Service