K8s Monitor
k8s-monitor is the primary data collection service. It polls the Kubernetes API and cloud provider pricing APIs on a configurable interval and exposes aggregated health, cost, and metrics data via a REST API.
What It Does
On every report cycle (executeReportCycle):
- Queries node and pod states from the Kubernetes API
- Computes a health score (0–100)
- Fetches cloud pricing data and calculates per-namespace costs
- Generates optimisation recommendations
- Appends a
MetricPointto the time-series store - Runs the statistical anomaly detector
- If anomalies are found and
RABBITMQ_URLis set, publishes tok8s.anomaliesexchange
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/health | Cluster health score, node status, control plane |
GET | /api/cost | Cost breakdown by namespace, workload, resource type |
GET | /api/optimizer | Over-provisioned and idle workload recommendations |
GET | /api/metrics/pods | Per-pod CPU and memory usage |
GET | /api/metrics/nodes | Per-node CPU, memory, and disk metrics |
GET | /api/history | Historical metric time-series (?start=&end=&cluster_id=) |
GET | /api/anomalies | Latest anomaly events from the statistical detector |
Health Score Formula
health_score = (
node_readiness_pct × 0.40 +
pod_success_rate × 0.30 +
control_plane_health × 0.20 +
api_latency_score × 0.10
) × 100
Values are clamped to 0–100. The score is returned as an integer in the health response alongside component breakdowns.
Time-Series Storage
k8s-monitor supports two storage backends, selected at startup:
- In-memory ring buffer (default) — goroutine-safe, holds the last N points (default 1000). Zero config. Data is lost on restart.
- PostgreSQL (opt-in via
DATABASE_URL) — persistsmetric_snapshotsto thek8s_monitorschema. Enables the/api/historyendpoint and feeds the predictive-scaler.
Anomaly Detection
The StatisticalDetector runs a Z-score calculation over a rolling window after each report cycle:
z = (current_value - rolling_mean) / rolling_std_dev
Default thresholds:
- Z-score ≥ 2.5 →
highseverity - Z-score ≥ 3.5 →
criticalseverity
Window size and thresholds are configurable via environment variables.
Environment Variables
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | Optional. Enables PostgreSQL time-series storage |
RABBITMQ_URL | — | Optional. Enables anomaly event publishing |
ANOMALY_WINDOW_SIZE | 60 | Number of data points in rolling window |
ANOMALY_Z_THRESHOLD | 2.5 | Z-score threshold for high severity |
CLUSTER_ID | default | Identifier attached to all metrics |
REPORT_INTERVAL | 30s | How often to run a health cycle |