K8s Optimizer
Repo: python/k8s-optimizer · Stack: Python 3.11+, scikit-learn, FastAPI
The only Python service in the platform. A RandomForest-based resource-rightsizing agent: it collects CPU/memory metrics from Prometheus and the Kubernetes API, trains an in-memory model per cluster, and recommends (or, outside dry-run, applies) scale-up/scale-down changes.
Key Capabilities
- Metrics collection —
MetricsCollectorpolls Prometheus and the Kubernetes API for per-workload resource usage - AI-driven predictions —
AIOptimizertrains a RandomForest model per resource dimension (CPU, memory) onceMIN_TRAINING_SAMPLES(default 100) samples are collected; retrains every 1,000 new samples. Models are not persisted — they re-warm from scratch on every process restart - Threshold-gated recommendations — a recommendation is only generated when predicted utilization crosses the configured scale-up/scale-down thresholds and the model's confidence exceeds
MIN_CONFIDENCE_SCORE(default0.7) - Dry-run mode —
DRY_RUN=true(or--dry-run) logs the patch body instead of calling the Kubernetes API; useful for validating recommendations before trusting the agent to act - Hard replica bounds —
MIN_REPLICAS/MAX_REPLICAScap what the optimizer will ever set, regardless of model output
Auth
None. Unlike the Go services in o-apps, this service has no auth middleware of any kind on its FastAPI routes — confirmed by reading api/server.py directly. Anyone who can reach it can read status/recommendations/feedback. Not yet part of the platform's auth-triage effort, since it lives outside the o-apps monorepo.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/status | Optimizer run status |
GET | /api/v1/recommendations | Current recommendations |
GET | /api/v1/feedback | Outcome feedback on previously-applied recommendations |
GET | /api/v1/metrics | Prometheus-format metrics export |
GET | /health | Health check |
Environment Variables
| Variable | Default | Purpose |
|---|---|---|
PROMETHEUS_URL | http://prometheus-server:9090 | Prometheus endpoint |
NAMESPACE | default | Target namespace |
OPTIMIZATION_INTERVAL | 300 | Seconds between optimization cycles |
CPU_SCALE_UP_THRESHOLD / CPU_SCALE_DOWN_THRESHOLD | 0.8 / 0.3 | CPU utilization ratios that trigger a recommendation |
MEMORY_SCALE_UP_THRESHOLD / MEMORY_SCALE_DOWN_THRESHOLD | 0.8 / 0.3 | Memory utilization ratios |
MIN_CONFIDENCE_SCORE | 0.7 | Minimum model confidence before a recommendation is generated |
MIN_TRAINING_SAMPLES | 100 | Samples required before the AI model activates |
MAX_REPLICAS / MIN_REPLICAS | 10 / 1 | Hard scaling bounds |
KUBECONFIG | — | Path to kubeconfig (omit for in-cluster) |
DRY_RUN | false | Log recommendations without applying them |