Observability Agent
Port: 8092 · DB Schema: observability
The observability agent is the first stage of the reactive AI pipeline. It collects telemetry on a configurable interval and publishes structured snapshots to the observability.telemetry RabbitMQ exchange.
What it collects, and from where
This service collects from more than just the host cluster. Every 5 minutes it refreshes its target list: a fixed set of host-cluster targets, plus one target per tenant's currently-Ready vCluster, discovered dynamically by logging in as a real Super-Admin-role system account and listing every CloudSpace and its vClusters via auth-service/kubeopera-api. A tenant's target is keyed by their vCluster's namespace — the same identifier the frontend's own tenant-namespace authorization already uses, so no separate cluster-ID-based authorization scheme was needed to keep this tenant-scoped. If discovery itself fails on a given cycle, the previous target list is kept rather than cleared, so a transient auth-service or kubeopera-api hiccup doesn't blank out telemetry collection entirely.
For each configured target, it collects from:
- k8s-monitor (
/api/health, or a namespace-scoped variant for tenant targets) — cluster/namespace health, node and pod counts, API latency, network throughput - security-api (
/api/v1/posture/summary) — host-cluster targets only, not per-tenant; security posture score
TelemetrySnapshot Fields
| Field | Unit | Description |
|---|---|---|
ClusterID | — | The target's identifier — a real cluster ID for host-cluster targets, or a tenant vCluster's namespace for tenant targets |
Source | — | k8s-monitor, security-api, or cicd-gateway |
HealthScore | 0–100 | Overall cluster/namespace health score |
CPUUsagePct | % | CPU utilisation |
MemoryUsagePct | % | Memory utilisation |
ReadyNodeCount | count | Nodes in Ready state |
NodeCount | count | All registered nodes |
PodCount | count | All pods observed |
FailedPodCount | count | Pods in Failed phase |
CrashLoopCount | count | Pods in CrashLoopBackOff |
APIServerLatencyMs | ms | API server response time |
NetworkRxBytesPS | bytes/s | Ingress traffic |
NetworkTxBytesPS | bytes/s | Egress traffic |
SecurityPostureScore | 0–100 | From security-api, host-cluster targets only |
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/telemetry/latest | Most recent snapshot per target |
GET | /api/v1/telemetry/history | Historical snapshots |
GET | /api/v1/telemetry/events | Telemetry-derived events |
GET | /api/v1/telemetry/clusters | Known cluster/target identifiers |
GET | /healthz | Health check |
Environment Variables
| Variable | Default | Description |
|---|---|---|
K8S_MONITOR_BASE_URL | http://k8s-monitor:8085 | k8s-monitor endpoint |
SECURITY_API_KEY | — | Shared secret for security-api's posture endpoint |
COLLECTION_INTERVAL | 30s | How often to collect telemetry |
DATABASE_URL | — | PostgreSQL connection |
RABBITMQ_URL | — | RabbitMQ connection |
AUTH_SERVICE_BASE_URL | — | Required for tenant target discovery |
KUBEOPERA_API_BASE_URL | — | Required for tenant target discovery |
AUTH_SERVICE_LOGIN / AUTH_SERVICE_PASSWORD | — | Credentials for the Super-Admin-role system account used to discover tenant vClusters. Tenant discovery is skipped (not an error) when these, along with the two URLs above, aren't all configured |
AUTH_APP_ID | KubeOpera's own app row | App ID used for the discovery login |
AUTH_JWT_ACCESS_SECRET | — | Validates bearer tokens on this service's own REST routes |
PORT | 8092 | HTTP port |