Skip to main content
Version: 1.0

K8s Monitor

k8s-monitor is the primary data collection service. It polls the Kubernetes API and cloud provider pricing APIs on a configurable interval and exposes aggregated health, cost, and metrics data via a REST API.

What It Does​

On every report cycle (executeReportCycle):

  1. Queries node and pod states from the Kubernetes API
  2. Computes a health score (0–100)
  3. Fetches cloud pricing data and calculates per-namespace costs
  4. Generates optimisation recommendations
  5. Appends a MetricPoint to the time-series store
  6. Runs the statistical anomaly detector
  7. If anomalies are found and RABBITMQ_URL is set, publishes to k8s.anomalies exchange

REST API​

MethodPathDescription
GET/api/healthCluster health score, node status, control plane
GET/api/costCost breakdown by namespace, workload, resource type
GET/api/optimizerOver-provisioned and idle workload recommendations
GET/api/metrics/podsPer-pod CPU and memory usage
GET/api/metrics/nodesPer-node CPU, memory, and disk metrics
GET/api/historyHistorical metric time-series (?start=&end=&cluster_id=)
GET/api/anomaliesLatest anomaly events from the statistical detector

Health Score Formula​

health_score = (
node_readiness_pct × 0.40 +
pod_success_rate × 0.30 +
control_plane_health × 0.20 +
api_latency_score × 0.10
) × 100

Values are clamped to 0–100. The score is returned as an integer in the health response alongside component breakdowns.

Time-Series Storage​

k8s-monitor supports two storage backends, selected at startup:

  • In-memory ring buffer (default) — goroutine-safe, holds the last N points (default 1000). Zero config. Data is lost on restart.
  • PostgreSQL (opt-in via DATABASE_URL) — persists metric_snapshots to the k8s_monitor schema. Enables the /api/history endpoint and feeds the predictive-scaler.

Anomaly Detection​

The StatisticalDetector runs a Z-score calculation over a rolling window after each report cycle:

z = (current_value - rolling_mean) / rolling_std_dev

Default thresholds:

  • Z-score ≥ 2.5 → high severity
  • Z-score ≥ 3.5 → critical severity

Window size and thresholds are configurable via environment variables.

Environment Variables​

VariableDefaultDescription
DATABASE_URL—Optional. Enables PostgreSQL time-series storage
RABBITMQ_URL—Optional. Enables anomaly event publishing
ANOMALY_WINDOW_SIZE60Number of data points in rolling window
ANOMALY_Z_THRESHOLD2.5Z-score threshold for high severity
CLUSTER_IDdefaultIdentifier attached to all metrics
REPORT_INTERVAL30sHow often to run a health cycle