Monitoring Guide
This guide covers two things:
- Monitoring your applications — how to instrument an app so KubeOpera's dashboards, SLOs, APM, tracing and AI agents can see inside it.
- Monitoring KubeOpera itself — how the platform's own services are scraped, visualized and alerted on.
Monitoring your applications
KubeOpera collects cluster-level signals — CPU, memory, restarts, pod status — for every workload automatically. To see inside your app (request rate, latency, errors, traces), give KubeOpera three things.
1. Expose Prometheus metrics
Serve metrics on a /metrics endpoint using a Prometheus client library for your language (Go, Python, Java, Node.js). At a minimum, record HTTP requests with the standard labels:
http_requests_total{service, route, method, status}
http_request_duration_seconds_bucket{service, route, method, le}
Then tell KubeOpera to scrape it. In the Deploy Application wizard, turn on Metrics and enter the port and path — or add the annotations yourself:
metadata:
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9898"
prometheus.io/path: "/metrics"
Your app's latency, throughput and error rate now appear in APM, on the Dashboard's observability panels and in the AI agents' view of your app.
2. Send traces
Instrument your app with OpenTelemetry and export traces with OTLP to KubeOpera's collector:
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector.monitoring:4317
OTEL_SERVICE_NAME=payments-api
Traces appear in Monitoring → Traces, and the service map shows how your services call each other.
3. Log to standard output
Write logs to stdout/stderr — ideally as JSON with level, msg and a trace_id. KubeOpera collects them with Loki automatically; they appear on your app's Logs tab, and log patterns feed root-cause analysis.
Define SLOs
Once your app exposes metrics, define service-level objectives in Monitoring → SLOs — for example 99.9% of checkout requests succeed over 30 days. KubeOpera tracks the error budget and burn rate and can alert you before the budget runs out. See slo-manager.
Monitoring KubeOpera itself
KubeOpera's services are monitored with the same tools as any workload: Prometheus, Grafana and Alertmanager, installed as part of every environment.
Metrics
Every KubeOpera service exposes Prometheus metrics on /metrics, on the same port as its API, and ships a ServiceMonitor so the Prometheus Operator scrapes it:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: kubeopera-api
spec:
selector:
matchLabels:
app: kubeopera-api
endpoints:
- port: http
interval: 30s
path: /metrics
Each service reports request rate, latency and errors per route, plus service-specific metrics — for example, messages consumed and processing time for event-driven services.
If a service's panels go flat, check that its ServiceMonitor port matches a port the container actually listens on. Prometheus's Targets page shows any target it can't scrape.
Dashboards
Grafana ships with dashboards for the platform, provisioned from Git alongside the services they cover:
- Platform overview — health of every KubeOpera service at a glance.
- Per-service — request rate, latency, errors and saturation.
- Reactive pipeline — telemetry throughput, analysis latency, actions taken and feedback signals.
- Message bus — RabbitMQ queue depth, publish and consume rates, and unroutable messages.
- Data stores — PostgreSQL and Redis.
Tenant-facing views — cost, performance, security, SLOs and traces for apps — are in the KubeOpera dashboard itself, under Monitoring.
Alerting
Alert rules are PrometheusRule resources in Git, evaluated by Prometheus and routed by Alertmanager. Out of the box they cover:
- elevated HTTP error rate or latency on any service;
- high CPU or memory per pod;
- growing RabbitMQ queues or unroutable messages;
- failed Flux reconciliations;
- certificates close to expiry.
Route alerts by severity to the channels your team uses — Slack, PagerDuty, Opsgenie or email — by editing the Alertmanager configuration in your environment:
route:
receiver: slack-alerts
routes:
- matchers: [ severity="critical" ]
receiver: pagerduty
receivers:
- name: slack-alerts
slack_configs:
- channel: "#alerts"
- name: pagerduty
pagerduty_configs:
- routing_key_file: /etc/alertmanager/secrets/pagerduty-key
Next steps
- Advanced monitoring — logs, traces, SLOs and APM programmatically.
- Dashboards reference — what each Monitoring page shows.
- Troubleshooting — what to check when something looks wrong.