Skip to main content
Version: 2.0

Predictive Scaler

Service: predictive-scaler · Port: 8089 · Database schema: forecasting

Autoscalers react after load arrives. predictive-scaler looks ahead: it forecasts each workload's load from its history and proposes scaling decisions so capacity is ready before demand peaks. You — or an authorized AI agent — approve them.

How forecasting works​

The scaler uses Holt's linear trend method (double exponential smoothing), implemented in Go with no ML framework dependencies. It tracks the metric's level and trend:

Level:    L(t) = α × y(t) + (1 − α) × (L(t−1) + T(t−1))
Trend: T(t) = β × (L(t) − L(t−1)) + (1 − β) × T(t−1)
Forecast: ŷ(t+h) = L(t) + h × T(t)
  • α (0–1) sets how quickly the level follows new data; β (0–1) does the same for the trend.
  • The prediction interval is ŷ ± 1.96 × RMSE — a 95% band.
  • A workload needs at least 24 data points before its first forecast.
  • Supported metrics: cpu_usage_pct, memory_usage_pct and request_rate.

For workloads with strong daily or weekly patterns, set FORECAST_SEASONALITY to use the seasonal (triple exponential smoothing) variant.

From forecast to decision​

When a forecast shows a workload's utilization will leave its comfortable range within the forecast horizon, the scaler proposes a scaling decision with the recommended replica count, the reason, and its confidence.

Approving a decision​

A decision can be approved in Analytics → Scaling Decisions, through the API, or by an AI agent with permission (the Cost Optimizer's approve_scaling_decision tool). On approval:

  1. If the workload has a HorizontalPodAutoscaler, the scaler raises the HPA's minReplicas to the recommended value, so the HPA and the forecast work together rather than against each other. Otherwise, it sets the Deployment's replicas directly.
  2. If the extra replicas need more nodes than are available, it asks nodes-manager to add capacity.
  3. It marks the decision applied, recording who approved it and when.

When the forecast peak has passed, the scaler proposes a follow-up decision to return to the original level.

Domain model​

Forecast
├── cluster_id, namespace, workload, metric
├── algorithm: holt | holt_winters
├── points: []{ timestamp, value, lower_bound, upper_bound }
├── confidence, horizon_mins
└── generated_at

ScalingDecision
├── id, forecast_id, cluster_id, namespace, workload
├── current_replicas, recommended_replicas
├── reason, confidence
├── status: pending | approved | applied | rejected
└── decided_by, executed_at

REST API​

MethodPathDescription
GET/api/v1/forecastsLatest forecasts (?cluster_id=&namespace=&workload=).
POST/api/v1/forecasts/generateForecast one workload now.
GET/api/v1/scaling-decisionsScaling decisions (?status=pending).
POST/api/v1/scaling-decisions/{id}/approveApprove and apply.
POST/api/v1/scaling-decisions/{id}/rejectReject ({ "reason": "…" }).
GET/healthzHealth check.

Permissions​

rules:
- apiGroups: ["autoscaling"]
resources: ["horizontalpodautoscalers"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["apps"]
resources: ["deployments", "deployments/scale"]
verbs: ["list", "get", "patch"]

Configuration​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection (reads k8s-monitor's metric history).
HOLT_WINTERS_ALPHA0.3Level smoothing factor.
HOLT_WINTERS_BETA0.1Trend smoothing factor.
FORECAST_SEASONALITY—Season length (for example 24h or 168h) to enable seasonal forecasting.
FORECAST_HORIZON_MINS60How far ahead to forecast.
NODES_MANAGER_BASE_URLhttp://nodes-manager:8115For adding node capacity.
PORT8089HTTP port.