Predictive Scaler
Service: predictive-scaler · Port: 8089 · Database schema: forecasting
Autoscalers react after load arrives. predictive-scaler looks ahead: it forecasts each workload's load from its history and proposes scaling decisions so capacity is ready before demand peaks. You — or an authorized AI agent — approve them.
How forecasting works
The scaler uses Holt's linear trend method (double exponential smoothing), implemented in Go with no ML framework dependencies. It tracks the metric's level and trend:
Level: L(t) = α × y(t) + (1 − α) × (L(t−1) + T(t−1))
Trend: T(t) = β × (L(t) − L(t−1)) + (1 − β) × T(t−1)
Forecast: ŷ(t+h) = L(t) + h × T(t)
- α (0–1) sets how quickly the level follows new data; β (0–1) does the same for the trend.
- The prediction interval is
ŷ ± 1.96 × RMSE— a 95% band. - A workload needs at least 24 data points before its first forecast.
- Supported metrics:
cpu_usage_pct,memory_usage_pctandrequest_rate.
For workloads with strong daily or weekly patterns, set FORECAST_SEASONALITY to use the seasonal (triple exponential smoothing) variant.
From forecast to decision
When a forecast shows a workload's utilization will leave its comfortable range within the forecast horizon, the scaler proposes a scaling decision with the recommended replica count, the reason, and its confidence.
Approving a decision
A decision can be approved in Analytics → Scaling Decisions, through the API, or by an AI agent with permission (the Cost Optimizer's approve_scaling_decision tool). On approval:
- If the workload has a HorizontalPodAutoscaler, the scaler raises the HPA's
minReplicasto the recommended value, so the HPA and the forecast work together rather than against each other. Otherwise, it sets the Deployment's replicas directly. - If the extra replicas need more nodes than are available, it asks nodes-manager to add capacity.
- It marks the decision
applied, recording who approved it and when.
When the forecast peak has passed, the scaler proposes a follow-up decision to return to the original level.
Domain model
Forecast
├── cluster_id, namespace, workload, metric
├── algorithm: holt | holt_winters
├── points: []{ timestamp, value, lower_bound, upper_bound }
├── confidence, horizon_mins
└── generated_at
ScalingDecision
├── id, forecast_id, cluster_id, namespace, workload
├── current_replicas, recommended_replicas
├── reason, confidence
├── status: pending | approved | applied | rejected
└── decided_by, executed_at
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/forecasts | Latest forecasts (?cluster_id=&namespace=&workload=). |
POST | /api/v1/forecasts/generate | Forecast one workload now. |
GET | /api/v1/scaling-decisions | Scaling decisions (?status=pending). |
POST | /api/v1/scaling-decisions/{id}/approve | Approve and apply. |
POST | /api/v1/scaling-decisions/{id}/reject | Reject ({ "reason": "…" }). |
GET | /healthz | Health check. |
Permissions
rules:
- apiGroups: ["autoscaling"]
resources: ["horizontalpodautoscalers"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["apps"]
resources: ["deployments", "deployments/scale"]
verbs: ["list", "get", "patch"]
Configuration
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection (reads k8s-monitor's metric history). |
HOLT_WINTERS_ALPHA | 0.3 | Level smoothing factor. |
HOLT_WINTERS_BETA | 0.1 | Trend smoothing factor. |
FORECAST_SEASONALITY | — | Season length (for example 24h or 168h) to enable seasonal forecasting. |
FORECAST_HORIZON_MINS | 60 | How far ahead to forecast. |
NODES_MANAGER_BASE_URL | http://nodes-manager:8115 | For adding node capacity. |
PORT | 8089 | HTTP port. |