Skip to main content
Version: 1.0

Predictive Autoscaler

predictive-scaler reads historical metric data from PostgreSQL, runs Holt-Winters exponential smoothing forecasts, and generates proactive scaling recommendations that operators or AI agents can approve.

Forecasting Algorithm​

Double exponential smoothing (Holt-Winters without seasonality) is implemented in pure Go — no Python or ML framework dependencies:

Level:   L(t) = α × y(t) + (1-α) × (L(t-1) + T(t-1))
Trend: T(t) = β × (L(t) - L(t-1)) + (1-β) × T(t-1)
Forecast at h steps ahead: ŷ(t+h) = L(t) + h × T(t)

Prediction interval: ŷ ± 1.96 × RMSE

Requirements:

  • Minimum 24 data points before a forecast is generated
  • Alpha (level smoothing) and beta (trend smoothing) are configurable
  • Supports metrics: cpu_usage_pct, memory_usage_pct, request_rate

Domain Model​

MetricSeries
└── cluster_id, namespace, workload, metric, value, timestamp

Forecast
├── cluster_id, namespace, workload, metric
├── algorithm: holt_winters | moving_avg
├── points: []ForecastPoint { timestamp, value, lower_bound, upper_bound }
├── confidence, horizon_mins
└── generated_at

ScalingDecision
├── forecast_id, cluster_id, namespace, workload
├── current_replicas, recommend_replicas
├── reason, confidence
├── status: pending | approved | applied | rejected
└── executed_at

REST API​

MethodPathDescription
GET/api/v1/forecastsList latest forecasts (?cluster_id=&namespace=&workload=)
POST/api/v1/forecasts/generateOn-demand forecast for a specific workload
GET/api/v1/scaling-decisionsList recommendations (?status=pending)
POST/api/v1/scaling-decisions/{id}/approveExecute scaling action
POST/api/v1/scaling-decisions/{id}/rejectReject recommendation

Approval Flow​

When a ScalingDecision is approved (via UI, API, or AI agent):

  1. predictive-scaler patches the workload's HPA spec.minReplicas to the recommended value
  2. If the recommendation requires new nodes (recommended replicas > current node capacity), it calls nodes-manager via HTTP to scale the node group
  3. Updates ScalingDecision.status to applied and records executed_at

Kubernetes RBAC Requirements​

rules:
- apiGroups: ["autoscaling"]
resources: ["horizontalpodautoscalers"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["apps"]
resources: ["deployments"]
verbs: ["list", "get", "patch"]

Environment Variables​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection (shared with k8s-monitor)
RABBITMQ_URL—Optional. For consuming metric snapshots via exchange
HOLT_WINTERS_ALPHA0.3Level smoothing factor (0–1)
HOLT_WINTERS_BETA0.1Trend smoothing factor (0–1)
FORECAST_HORIZON_MINS60How far ahead to forecast
NODES_MANAGER_BASE_URL—URL for node group scaling
PORT8089HTTP port