K8s Monitor
Service: k8s-monitor · Port: 8085 · Database schema: k8s_monitor
k8s-monitor is KubeOpera's eyes on your clusters. On a regular cycle it reads the state of every registered cluster from the Kubernetes API and cloud pricing APIs, and turns it into the numbers you see everywhere in KubeOpera: health scores, cost, metrics and right-sizing recommendations. It also watches for statistical anomalies and publishes them for the automation services.
What happens each cycle
Every 30 seconds (by default), for each cluster:
- Read state — nodes, pods, control-plane components and API latency.
- Score health — compute a 0–100 health score (see below).
- Price it — fetch current pricing and compute cost per node, namespace and workload.
- Find waste — flag over-provisioned and idle workloads.
- Record — store a
MetricPointin the time-series history. - Detect anomalies — compare each metric with its rolling baseline.
- Publish — send any anomalies to the
k8s.anomaliesexchange.
Health score
The score starts at 100 and loses points for each problem found:
| Condition | Penalty |
|---|---|
| Each not-ready node | −5 |
| Each node under memory pressure | −3 |
| Each node under disk pressure | −4 |
| Each node under PID pressure | −2 |
| Each node with network unavailable | −6 |
| Each failed pod | −2 |
| Each crash-looping pod | −2.5 |
| More than 10 restarting pods | −5 |
| Control plane degraded | −30 (−15 if only the API server is affected) |
| Each unhealthy control-plane component | −8 |
| Cluster CPU ≥ 95% (≥ 80%: −5) | −10 |
| Cluster memory ≥ 95% (≥ 80%: −5) | −10 |
The result is kept between 0 and 100. The health response includes every issue that cost points, with a message and — where possible — a suggested fix, so you can see exactly why a score is what it is.
Cost
k8s-monitor prices each node from its cloud provider's on-demand or spot price (AWS and GCP), or from a configurable rate card for on-premises nodes. It allocates node cost to pods by their resource requests, then rolls it up by namespace, workload and tenant.
History
Metric points are stored in PostgreSQL, so history survives restarts, charts can show long time ranges, and the predictive scaler has data to forecast from.
Anomaly detection
After each cycle, a statistical detector compares each metric with its rolling baseline:
z = (current_value − rolling_mean) / rolling_std_dev
A Z-score of 2.5 or more is high severity; 3.5 or more is critical. Anomalies are published to k8s.anomalies with the routing key anomaly.k8s-monitor.{severity} for the anomaly detector and incident manager.
This detector feeds rule-based remediation. The adaptive, learning detector used by the reactive AI pipeline runs in the analysis agent.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/health | Health score, issues, node and control-plane status. |
GET | /api/health/namespace/{ns} | Health for one namespace or tenant vCluster. |
GET | /api/cost | Cost by namespace, workload and resource type. |
GET | /api/optimizer | Over-provisioned and idle workloads, with suggested sizes. |
GET | /api/metrics/pods | CPU and memory per pod (?namespace=). |
GET | /api/metrics/nodes | CPU, memory and disk per node. |
GET | /api/history | Metric history (?cluster_id=&start=&end=). |
GET | /api/anomalies | Recent anomalies. |
GET | /api/combined | Health, cost and metrics in one call (used by the Dashboard). |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
REPORT_INTERVAL | 30s | How often to run a cycle. |
DATABASE_URL | — | PostgreSQL connection (required). |
RABBITMQ_URL | — | RabbitMQ connection, for anomaly publishing. |
ANOMALY_WINDOW_SIZE | 60 | Data points in the rolling window. |
ANOMALY_Z_THRESHOLD_HIGH | 2.5 | Z-score for high severity. |
ANOMALY_Z_THRESHOLD_CRITICAL | 3.5 | Z-score for critical severity. |
ONPREM_RATE_CARD | — | Path to a rate card for pricing on-premises nodes. |
AUTH_JWT_ACCESS_SECRET | — | Validates bearer tokens. |
PORT | 8085 | HTTP port. |