Skip to main content
Version: 2.0

K8s Optimizer

Service: k8s-optimizer · Port: 8097 · Stack: Python, scikit-learn, FastAPI

k8s-optimizer uses machine learning to right-size workloads. It learns each workload's CPU and memory patterns, predicts where utilization is heading, and recommends — or, if you allow it, applies — replica changes before a workload becomes over- or under-provisioned.

How it works​

  1. Collect — gather per-workload CPU and memory usage from Prometheus and the Kubernetes API.
  2. Learn — train a RandomForest model per workload and resource once it has enough history (100 samples by default), and retrain as new data arrives. Models are saved, so they survive restarts.
  3. Predict — forecast each workload's utilization.
  4. Recommend — propose a scale-up or scale-down when predicted utilization crosses your thresholds and the model is confident enough (0.7 by default).
  5. Apply (optional) — outside dry-run mode, apply the change within your replica bounds.
  6. Learn from outcomes — record whether each change brought utilization back into range, and use that feedback in future predictions.

Safety​

  • Dry-run mode (DRY_RUN=true) records recommendations without changing anything — start here to build trust.
  • Hard bounds (MIN_REPLICAS, MAX_REPLICAS) cap every change, whatever the model says.
  • Confidence gating — low-confidence predictions never produce recommendations.
  • Applied changes go through the action agent, so they respect PodDisruptionBudgets and are audited.

REST API​

All routes require a valid KubeOpera token.

MethodPathDescription
GET/api/v1/statusOptimizer status and model readiness per workload.
GET/api/v1/recommendationsCurrent recommendations with confidence.
GET/api/v1/feedbackOutcomes of previously applied recommendations.
GET/metricsPrometheus metrics.
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
PROMETHEUS_URLhttp://kube-prometheus-stack-prometheus.monitoring:9090Prometheus.
NAMESPACESallNamespaces to optimize (comma-separated).
OPTIMIZATION_INTERVAL300Seconds between cycles.
CPU_SCALE_UP_THRESHOLD / CPU_SCALE_DOWN_THRESHOLD0.8 / 0.3CPU utilization that triggers a recommendation.
MEMORY_SCALE_UP_THRESHOLD / MEMORY_SCALE_DOWN_THRESHOLD0.8 / 0.3Memory utilization that triggers a recommendation.
MIN_CONFIDENCE_SCORE0.7Minimum model confidence.
MIN_TRAINING_SAMPLES100History needed before predictions start.
MIN_REPLICAS / MAX_REPLICAS1 / 10Replica bounds.
MODEL_STORE_PATH/data/modelsWhere trained models are saved.
DRY_RUNtrueRecommend without applying.
AUTH_JWT_ACCESS_SECRET—Validates tokens.
PORT8097HTTP port.