Skip to main content
Version: 2.0

Monitoring Guide

This guide covers two things:

  1. Monitoring your applications — how to instrument an app so KubeOpera's dashboards, SLOs, APM, tracing and AI agents can see inside it.
  2. Monitoring KubeOpera itself — how the platform's own services are scraped, visualized and alerted on.

Monitoring your applications​

KubeOpera collects cluster-level signals — CPU, memory, restarts, pod status — for every workload automatically. To see inside your app (request rate, latency, errors, traces), give KubeOpera three things.

1. Expose Prometheus metrics​

Serve metrics on a /metrics endpoint using a Prometheus client library for your language (Go, Python, Java, Node.js). At a minimum, record HTTP requests with the standard labels:

http_requests_total{service, route, method, status}
http_request_duration_seconds_bucket{service, route, method, le}

Then tell KubeOpera to scrape it. In the Deploy Application wizard, turn on Metrics and enter the port and path — or add the annotations yourself:

metadata:
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "9898"
prometheus.io/path: "/metrics"

Your app's latency, throughput and error rate now appear in APM, on the Dashboard's observability panels and in the AI agents' view of your app.

2. Send traces​

Instrument your app with OpenTelemetry and export traces with OTLP to KubeOpera's collector:

OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector.monitoring:4317
OTEL_SERVICE_NAME=payments-api

Traces appear in Monitoring → Traces, and the service map shows how your services call each other.

3. Log to standard output​

Write logs to stdout/stderr — ideally as JSON with level, msg and a trace_id. KubeOpera collects them with Loki automatically; they appear on your app's Logs tab, and log patterns feed root-cause analysis.

Define SLOs​

Once your app exposes metrics, define service-level objectives in Monitoring → SLOs — for example 99.9% of checkout requests succeed over 30 days. KubeOpera tracks the error budget and burn rate and can alert you before the budget runs out. See slo-manager.

Monitoring KubeOpera itself​

KubeOpera's services are monitored with the same tools as any workload: Prometheus, Grafana and Alertmanager, installed as part of every environment.

Metrics​

Every KubeOpera service exposes Prometheus metrics on /metrics, on the same port as its API, and ships a ServiceMonitor so the Prometheus Operator scrapes it:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: kubeopera-api
spec:
selector:
matchLabels:
app: kubeopera-api
endpoints:
- port: http
interval: 30s
path: /metrics

Each service reports request rate, latency and errors per route, plus service-specific metrics — for example, messages consumed and processing time for event-driven services.

tip

If a service's panels go flat, check that its ServiceMonitor port matches a port the container actually listens on. Prometheus's Targets page shows any target it can't scrape.

Dashboards​

Grafana ships with dashboards for the platform, provisioned from Git alongside the services they cover:

  • Platform overview — health of every KubeOpera service at a glance.
  • Per-service — request rate, latency, errors and saturation.
  • Reactive pipeline — telemetry throughput, analysis latency, actions taken and feedback signals.
  • Message bus — RabbitMQ queue depth, publish and consume rates, and unroutable messages.
  • Data stores — PostgreSQL and Redis.

Tenant-facing views — cost, performance, security, SLOs and traces for apps — are in the KubeOpera dashboard itself, under Monitoring.

Alerting​

Alert rules are PrometheusRule resources in Git, evaluated by Prometheus and routed by Alertmanager. Out of the box they cover:

  • elevated HTTP error rate or latency on any service;
  • high CPU or memory per pod;
  • growing RabbitMQ queues or unroutable messages;
  • failed Flux reconciliations;
  • certificates close to expiry.

Route alerts by severity to the channels your team uses — Slack, PagerDuty, Opsgenie or email — by editing the Alertmanager configuration in your environment:

route:
receiver: slack-alerts
routes:
- matchers: [ severity="critical" ]
receiver: pagerduty
receivers:
- name: slack-alerts
slack_configs:
- channel: "#alerts"
- name: pagerduty
pagerduty_configs:
- routing_key_file: /etc/alertmanager/secrets/pagerduty-key

Next steps​