Skip to main content
Version: 1.0

Analytics

The Analytics page (/analytics) consolidates predictive and anomaly data from predictive-scaler and anomaly-detector into three tabs: Forecasts, Anomalies, and Scaling Decisions.

Summary cards​

Four cards above the tabs: Active Forecasts (count), Critical Anomalies, Pending Decisions, and Total Anomalies — all counted directly from what's currently loaded, not separately fetched aggregates.

Forecasts Tab​

Displays a forecast card per workload, using recharts' ComposedChart — a shaded confidence-band area plus a mean-forecast line. There's no workload/time-range filter UI on this page; every returned forecast is simply rendered as its own card.

How Forecasts Work​

predictive-scaler reads historical metric series and runs Holt's linear trend method (the service calls it "Holt-Winters," though there's no seasonal component — it's level + trend only, not the full triple-exponential-smoothing algorithm):

Level:    L(t) = α·y(t) + (1-α)·(L(t-1) + T(t-1))
Trend: T(t) = β·(L(t) - L(t-1)) + (1-β)·T(t-1)
Forecast: L(t) + h·T(t) (h steps ahead)

Alpha and beta are hardcoded at 0.3 and 0.1 respectively, not configurable via environment variables — every forecast uses the same smoothing weights. A minimum of 24 data points is required before a forecast is generated; below that, the service returns an "insufficient data" error rather than a low-confidence guess. The confidence band shown is ±1.96×RMSE.

Anomalies Tab​

A table of anomaly events from anomaly-detector.

Column shownDescription
Metrice.g. cpu_usage_pct, crash_loop_count
ResourceThe affected resource or namespace
Value / Z-scoreObserved value, with the Z-score in parentheses
Severitylow / medium / high / critical
DetectedRelative time since detection

There's no filter UI (cluster/severity/date range) on this page today — the table shows whatever the last 50 anomalies query returns. Two more fields exist on every anomaly event but aren't rendered as table columns here: baseline (the rolling-mean value the anomaly was measured against) and acknowledged (a boolean, flippable via a dedicated POST /api/v1/anomalies/{id}/acknowledge endpoint on anomaly-detector). There's no boolean "healing" flag on an event — instead, an event that triggered self-healing carries a healing_id referencing a separate remediation-log record, which has its own pending/executing/success/failed status.

Scaling Decisions Tab​

Lists ScalingDecision records from predictive-scaler. Each card shows:

  • Workload (namespace/name)
  • Current → recommended replicas, with a scale-up/down indicator
  • Reason and confidence percentage
  • Status: pending / approved / applied / rejected

Approve (POST /api/forecasts/api/v1/scaling-decisions/{id}/approve) patches the target Deployment's scale subresource directly — it sets spec.replicas via the Kubernetes API, not an HPA's minReplicas. There's no HPA involved anywhere in this flow. Reject (POST /api/forecasts/api/v1/scaling-decisions/{id}/reject) just marks the decision rejected in the database; it never touches the cluster.

Approving or rejecting a scaling decision is a human action today — there's no AI agent tool that can do either on your behalf. agent-runtime's Cost Optimizer agent can read forecasts (get_scaling_forecasts) but has no write/approve capability over them.