K8s Optimizer
Service: k8s-optimizer · Port: 8097 · Stack: Python, scikit-learn, FastAPI
k8s-optimizer uses machine learning to right-size workloads. It learns each workload's CPU and memory patterns, predicts where utilization is heading, and recommends — or, if you allow it, applies — replica changes before a workload becomes over- or under-provisioned.
How it works
- Collect — gather per-workload CPU and memory usage from Prometheus and the Kubernetes API.
- Learn — train a RandomForest model per workload and resource once it has enough history (100 samples by default), and retrain as new data arrives. Models are saved, so they survive restarts.
- Predict — forecast each workload's utilization.
- Recommend — propose a scale-up or scale-down when predicted utilization crosses your thresholds and the model is confident enough (0.7 by default).
- Apply (optional) — outside dry-run mode, apply the change within your replica bounds.
- Learn from outcomes — record whether each change brought utilization back into range, and use that feedback in future predictions.
Safety
- Dry-run mode (
DRY_RUN=true) records recommendations without changing anything — start here to build trust. - Hard bounds (
MIN_REPLICAS,MAX_REPLICAS) cap every change, whatever the model says. - Confidence gating — low-confidence predictions never produce recommendations.
- Applied changes go through the action agent, so they respect PodDisruptionBudgets and are audited.
REST API
All routes require a valid KubeOpera token.
| Method | Path | Description |
|---|---|---|
GET | /api/v1/status | Optimizer status and model readiness per workload. |
GET | /api/v1/recommendations | Current recommendations with confidence. |
GET | /api/v1/feedback | Outcomes of previously applied recommendations. |
GET | /metrics | Prometheus metrics. |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
PROMETHEUS_URL | http://kube-prometheus-stack-prometheus.monitoring:9090 | Prometheus. |
NAMESPACES | all | Namespaces to optimize (comma-separated). |
OPTIMIZATION_INTERVAL | 300 | Seconds between cycles. |
CPU_SCALE_UP_THRESHOLD / CPU_SCALE_DOWN_THRESHOLD | 0.8 / 0.3 | CPU utilization that triggers a recommendation. |
MEMORY_SCALE_UP_THRESHOLD / MEMORY_SCALE_DOWN_THRESHOLD | 0.8 / 0.3 | Memory utilization that triggers a recommendation. |
MIN_CONFIDENCE_SCORE | 0.7 | Minimum model confidence. |
MIN_TRAINING_SAMPLES | 100 | History needed before predictions start. |
MIN_REPLICAS / MAX_REPLICAS | 1 / 10 | Replica bounds. |
MODEL_STORE_PATH | /data/models | Where trained models are saved. |
DRY_RUN | true | Recommend without applying. |
AUTH_JWT_ACCESS_SECRET | — | Validates tokens. |
PORT | 8097 | HTTP port. |