Skip to main content
Version: 1.0

K8s Optimizer

Repo: python/k8s-optimizer · Stack: Python 3.11+, scikit-learn, FastAPI

The only Python service in the platform. A RandomForest-based resource-rightsizing agent: it collects CPU/memory metrics from Prometheus and the Kubernetes API, trains an in-memory model per cluster, and recommends (or, outside dry-run, applies) scale-up/scale-down changes.

Key Capabilities​

  • Metrics collection — MetricsCollector polls Prometheus and the Kubernetes API for per-workload resource usage
  • AI-driven predictions — AIOptimizer trains a RandomForest model per resource dimension (CPU, memory) once MIN_TRAINING_SAMPLES (default 100) samples are collected; retrains every 1,000 new samples. Models are not persisted — they re-warm from scratch on every process restart
  • Threshold-gated recommendations — a recommendation is only generated when predicted utilization crosses the configured scale-up/scale-down thresholds and the model's confidence exceeds MIN_CONFIDENCE_SCORE (default 0.7)
  • Dry-run mode — DRY_RUN=true (or --dry-run) logs the patch body instead of calling the Kubernetes API; useful for validating recommendations before trusting the agent to act
  • Hard replica bounds — MIN_REPLICAS/MAX_REPLICAS cap what the optimizer will ever set, regardless of model output

Auth​

None. Unlike the Go services in o-apps, this service has no auth middleware of any kind on its FastAPI routes — confirmed by reading api/server.py directly. Anyone who can reach it can read status/recommendations/feedback. Not yet part of the platform's auth-triage effort, since it lives outside the o-apps monorepo.

REST API​

MethodPathDescription
GET/api/v1/statusOptimizer run status
GET/api/v1/recommendationsCurrent recommendations
GET/api/v1/feedbackOutcome feedback on previously-applied recommendations
GET/api/v1/metricsPrometheus-format metrics export
GET/healthHealth check

Environment Variables​

VariableDefaultPurpose
PROMETHEUS_URLhttp://prometheus-server:9090Prometheus endpoint
NAMESPACEdefaultTarget namespace
OPTIMIZATION_INTERVAL300Seconds between optimization cycles
CPU_SCALE_UP_THRESHOLD / CPU_SCALE_DOWN_THRESHOLD0.8 / 0.3CPU utilization ratios that trigger a recommendation
MEMORY_SCALE_UP_THRESHOLD / MEMORY_SCALE_DOWN_THRESHOLD0.8 / 0.3Memory utilization ratios
MIN_CONFIDENCE_SCORE0.7Minimum model confidence before a recommendation is generated
MIN_TRAINING_SAMPLES100Samples required before the AI model activates
MAX_REPLICAS / MIN_REPLICAS10 / 1Hard scaling bounds
KUBECONFIG—Path to kubeconfig (omit for in-cluster)
DRY_RUNfalseLog recommendations without applying them