Skip to main content
Version: 1.0

Monitoring Guide

KubeOpera's own operation is monitored the same way any Kubernetes workload on the platform is — Prometheus scraping, Grafana dashboards, and AlertManager routing — rather than through a separate monitoring subsystem or CLI. This page covers how that's wired up for the platform's own services; monitoring a tenant's apps is a different, dashboard-driven experience covered in Platform Overview and the various app/monitoring/* pages it links to.

How services get scraped​

Each backend service that exposes Prometheus metrics carries its own ServiceMonitor custom resource, which tells the cluster's Prometheus Operator where to scrape it from and how often. kubeopera-api's, for example, points at its http port (the same port the service's main API listens on — Go's net/http only ever binds once, so /metrics lives on that router rather than a separate port) on a 30-second interval:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: kubeopera-backend
spec:
selector:
matchLabels:
app: kubeopera-backend
endpoints:
- port: http
interval: 30s
path: /metrics

This is worth calling out specifically because it's a mistake easy to make and easy to leave unnoticed: a ServiceMonitor pointing at a port nothing is actually listening on scrapes successfully as far as Prometheus is concerned (it just gets connection errors, which don't always surface loudly), so metrics silently stop flowing rather than failing in an obvious way. If a service's dashboard panels go flat, checking that the ServiceMonitor's port actually matches something the container listens on is worth doing before assuming the service itself is broken.

Dashboards​

Grafana is deployed as part of the platform's own monitoring stack, with dashboards provisioned alongside the services they cover rather than built ad hoc in the UI. If you're looking for a tenant-facing view of cost, performance, security posture, SLOs, or traces instead of the platform's own internals, those live in the dashboard itself under Monitoring, not in Grafana.

Alerting​

Alert rules are PrometheusRule resources, evaluated continuously and routed through AlertManager to Slack. The rules that exist today cover the fundamentals — elevated HTTP error rate, elevated latency, and high per-pod CPU or memory usage — routed by severity: routine alerts land in #alerts, warnings in #warnings, and anything critical in #critical-alerts. There's no PagerDuty integration configured today; if your team needs one, it's a real gap worth raising rather than something to assume exists because a lot of monitoring stacks have it.

Next Steps​