Skip to main content
Version: 2.0

K8s Monitor

Service: k8s-monitor · Port: 8085 · Database schema: k8s_monitor

k8s-monitor is KubeOpera's eyes on your clusters. On a regular cycle it reads the state of every registered cluster from the Kubernetes API and cloud pricing APIs, and turns it into the numbers you see everywhere in KubeOpera: health scores, cost, metrics and right-sizing recommendations. It also watches for statistical anomalies and publishes them for the automation services.

What happens each cycle​

Every 30 seconds (by default), for each cluster:

  1. Read state — nodes, pods, control-plane components and API latency.
  2. Score health — compute a 0–100 health score (see below).
  3. Price it — fetch current pricing and compute cost per node, namespace and workload.
  4. Find waste — flag over-provisioned and idle workloads.
  5. Record — store a MetricPoint in the time-series history.
  6. Detect anomalies — compare each metric with its rolling baseline.
  7. Publish — send any anomalies to the k8s.anomalies exchange.

Health score​

The score starts at 100 and loses points for each problem found:

ConditionPenalty
Each not-ready node−5
Each node under memory pressure−3
Each node under disk pressure−4
Each node under PID pressure−2
Each node with network unavailable−6
Each failed pod−2
Each crash-looping pod−2.5
More than 10 restarting pods−5
Control plane degraded−30 (−15 if only the API server is affected)
Each unhealthy control-plane component−8
Cluster CPU ≥ 95% (≥ 80%: −5)−10
Cluster memory ≥ 95% (≥ 80%: −5)−10

The result is kept between 0 and 100. The health response includes every issue that cost points, with a message and — where possible — a suggested fix, so you can see exactly why a score is what it is.

Cost​

k8s-monitor prices each node from its cloud provider's on-demand or spot price (AWS and GCP), or from a configurable rate card for on-premises nodes. It allocates node cost to pods by their resource requests, then rolls it up by namespace, workload and tenant.

History​

Metric points are stored in PostgreSQL, so history survives restarts, charts can show long time ranges, and the predictive scaler has data to forecast from.

Anomaly detection​

After each cycle, a statistical detector compares each metric with its rolling baseline:

z = (current_value − rolling_mean) / rolling_std_dev

A Z-score of 2.5 or more is high severity; 3.5 or more is critical. Anomalies are published to k8s.anomalies with the routing key anomaly.k8s-monitor.{severity} for the anomaly detector and incident manager.

This detector feeds rule-based remediation. The adaptive, learning detector used by the reactive AI pipeline runs in the analysis agent.

REST API​

MethodPathDescription
GET/api/healthHealth score, issues, node and control-plane status.
GET/api/health/namespace/{ns}Health for one namespace or tenant vCluster.
GET/api/costCost by namespace, workload and resource type.
GET/api/optimizerOver-provisioned and idle workloads, with suggested sizes.
GET/api/metrics/podsCPU and memory per pod (?namespace=).
GET/api/metrics/nodesCPU, memory and disk per node.
GET/api/historyMetric history (?cluster_id=&start=&end=).
GET/api/anomaliesRecent anomalies.
GET/api/combinedHealth, cost and metrics in one call (used by the Dashboard).
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
REPORT_INTERVAL30sHow often to run a cycle.
DATABASE_URL—PostgreSQL connection (required).
RABBITMQ_URL—RabbitMQ connection, for anomaly publishing.
ANOMALY_WINDOW_SIZE60Data points in the rolling window.
ANOMALY_Z_THRESHOLD_HIGH2.5Z-score for high severity.
ANOMALY_Z_THRESHOLD_CRITICAL3.5Z-score for critical severity.
ONPREM_RATE_CARD—Path to a rate card for pricing on-premises nodes.
AUTH_JWT_ACCESS_SECRET—Validates bearer tokens.
PORT8085HTTP port.