Skip to main content
Version: 2.0

Optimizer

Service: kubeopera-ai · Port: 8113

The Optimizer is an AI operations assistant that watches your cluster continuously. It gives you:

  • a regular, AI-written health assessment with ranked issues and fixes;
  • resource optimization reports with ready-to-apply changes and estimated savings;
  • load-test analysis that finds bottlenecks and recommends scaling;
  • an interactive troubleshooter that investigates using kubectl and Prometheus.

Results appear on the AI Optimizer dashboard page and in agent runs.

Continuous health assessment​

Every cycle (five minutes by default), the Optimizer gathers node status, pod states, recent events and resource usage, and asks Claude for a structured assessment:

  • a health score from 0 to 100 and an overall status;
  • issues, ranked by severity, each with a description, a recommendation and — where one applies — the exact command to fix it.

Critical and warning issues are published to k8s.anomalies (routing key anomaly.kubeopera-ai.{severity}), where the anomaly detector and incident manager pick them up — so an issue the Optimizer spots can open an incident, trigger remediation, and raise the risk that launches an SRE investigation.

{
"cluster_id": "prod-us-east",
"source": "kubeopera-ai",
"metric": "payments-api",
"value": 32.0,
"z_score": 3.0,
"baseline": 68.0,
"severity": "critical",
"message": "Multiple pods in CrashLoopBackOff, memory pressure detected"
}

Resource optimization​

The Optimizer compares seven days of CPU and memory usage (from apm-gateway) with each workload's requests and limits, and produces a report with:

  • right-sized requests and limits for each workload;
  • ready-to-run kubectl patch commands (or one-click apply from the dashboard, through GitOps);
  • the estimated monthly saving.

Load-test analysis​

Give it a load test's results and it returns a p95/p99 assessment, likely bottlenecks, scaling recommendations and a prioritized action list — comparing against its own baseline and live metrics. The Load Test Analyst agent builds on this.

Interactive troubleshooter​

POST /api/v1/troubleshoot starts a conversation in which Claude investigates your question with two narrowly-scoped tools:

  • execute_kubectl — read-only kubectl commands, run without a shell (arguments are passed directly, so there's no injection risk);
  • query_prometheus — PromQL queries through apm-gateway.

Conversations are bounded (ten tool iterations, with older history summarized) so they stay fast and focused.

Integration with agents​

Two tools expose the Optimizer to every agent:

  • get_optimization_status — the latest health assessment and optimization report;
  • get_load_test_analysis — run a load-test analysis.

REST API​

MethodPathDescription
GET/api/v1/statusLatest health assessment, optimization report and uptime.
POST/api/v1/analyze/healthAssess health now.
POST/api/v1/analyze/optimizeProduce an optimization report now.
POST/api/v1/analyze/load-testAnalyze load-test results ({ "results_path": "…" }).
POST/api/v1/troubleshootAsk the troubleshooter ({ "query": "…" }).
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
INTERVAL5mTime between health assessments.
NAMESPACE—Namespace to focus on (all namespaces when unset).
APM_GATEWAY_BASE_URLhttp://apm-gateway:8101Source of usage history and PromQL.
RABBITMQ_URL—Publishes issues to k8s.anomalies.
AUTH_JWT_ACCESS_SECRET—Validates tokens (required).
AUTH_SERVICE_BASE_URL / AI_CREDENTIAL_INTERNAL_API_KEY—AI credential resolution.
PORT8113HTTP port.