Optimizer
Service: kubeopera-ai · Port: 8113
The Optimizer is an AI operations assistant that watches your cluster continuously. It gives you:
- a regular, AI-written health assessment with ranked issues and fixes;
- resource optimization reports with ready-to-apply changes and estimated savings;
- load-test analysis that finds bottlenecks and recommends scaling;
- an interactive troubleshooter that investigates using kubectl and Prometheus.
Results appear on the AI Optimizer dashboard page and in agent runs.
Continuous health assessment
Every cycle (five minutes by default), the Optimizer gathers node status, pod states, recent events and resource usage, and asks Claude for a structured assessment:
- a health score from 0 to 100 and an overall status;
- issues, ranked by severity, each with a description, a recommendation and — where one applies — the exact command to fix it.
Critical and warning issues are published to k8s.anomalies (routing key anomaly.kubeopera-ai.{severity}), where the anomaly detector and incident manager pick them up — so an issue the Optimizer spots can open an incident, trigger remediation, and raise the risk that launches an SRE investigation.
{
"cluster_id": "prod-us-east",
"source": "kubeopera-ai",
"metric": "payments-api",
"value": 32.0,
"z_score": 3.0,
"baseline": 68.0,
"severity": "critical",
"message": "Multiple pods in CrashLoopBackOff, memory pressure detected"
}
Resource optimization
The Optimizer compares seven days of CPU and memory usage (from apm-gateway) with each workload's requests and limits, and produces a report with:
- right-sized requests and limits for each workload;
- ready-to-run
kubectl patchcommands (or one-click apply from the dashboard, through GitOps); - the estimated monthly saving.
Load-test analysis
Give it a load test's results and it returns a p95/p99 assessment, likely bottlenecks, scaling recommendations and a prioritized action list — comparing against its own baseline and live metrics. The Load Test Analyst agent builds on this.
Interactive troubleshooter
POST /api/v1/troubleshoot starts a conversation in which Claude investigates your question with two narrowly-scoped tools:
execute_kubectl— read-only kubectl commands, run without a shell (arguments are passed directly, so there's no injection risk);query_prometheus— PromQL queries through apm-gateway.
Conversations are bounded (ten tool iterations, with older history summarized) so they stay fast and focused.
Integration with agents
Two tools expose the Optimizer to every agent:
get_optimization_status— the latest health assessment and optimization report;get_load_test_analysis— run a load-test analysis.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/status | Latest health assessment, optimization report and uptime. |
POST | /api/v1/analyze/health | Assess health now. |
POST | /api/v1/analyze/optimize | Produce an optimization report now. |
POST | /api/v1/analyze/load-test | Analyze load-test results ({ "results_path": "…" }). |
POST | /api/v1/troubleshoot | Ask the troubleshooter ({ "query": "…" }). |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
INTERVAL | 5m | Time between health assessments. |
NAMESPACE | — | Namespace to focus on (all namespaces when unset). |
APM_GATEWAY_BASE_URL | http://apm-gateway:8101 | Source of usage history and PromQL. |
RABBITMQ_URL | — | Publishes issues to k8s.anomalies. |
AUTH_JWT_ACCESS_SECRET | — | Validates tokens (required). |
AUTH_SERVICE_BASE_URL / AI_CREDENTIAL_INTERNAL_API_KEY | — | AI credential resolution. |
PORT | 8113 | HTTP port. |