Observability Agent
Service: observability-agent-srv · Port: 8092 · Database schema: observability
The observability agent is the first stage of the reactive pipeline. On every cycle it collects telemetry for each cluster and tenant, stores it as a snapshot, and publishes it to the observability.telemetry exchange for the analysis agent. Its history also powers the charts and the /agents/observability page.
What it collects
Targets
The agent collects from two kinds of target:
- Host clusters — every cluster registered in KubeOpera.
- Tenant vClusters — every
ReadyvCluster in every CloudSpace, identified by its vCluster namespace.
It refreshes the target list every five minutes by asking kubeopera-api for registered clusters and CloudSpaces, using its own service identity with read-only platform scope. If a refresh fails, the previous target list is kept, so a brief outage elsewhere never interrupts collection.
Sources
For each target it gathers:
| Source | What it provides |
|---|---|
| k8s-monitor | Health score, node and pod counts, failures and crash loops, API server latency, network throughput. Tenant targets use a namespace-scoped view. |
| security-api | The security posture score — for host clusters and for each tenant's own workloads. |
| cicd-gateway | Recent deployments (from cicd.events), so later stages can relate anomalies to changes. |
Telemetry snapshot
| Field | Unit | Description |
|---|---|---|
ClusterID | — | The target: a host cluster ID or a tenant vCluster namespace. |
Source | — | k8s-monitor, security-api or cicd-gateway. |
HealthScore | 0–100 | Overall health of the cluster or tenant. |
CPUUsagePct | % | CPU utilization. |
MemoryUsagePct | % | Memory utilization. |
ReadyNodeCount | count | Nodes in the Ready state. |
NodeCount | count | All nodes. |
PodCount | count | All pods observed. |
FailedPodCount | count | Pods in the Failed phase. |
CrashLoopCount | count | Pods in CrashLoopBackOff. |
APIServerLatencyMs | ms | API server response time. |
NetworkRxBytesPS | bytes/s | Inbound traffic. |
NetworkTxBytesPS | bytes/s | Outbound traffic. |
SecurityPostureScore | 0–100 | Security posture from security-api. |
Events
| Direction | Exchange | Message |
|---|---|---|
| Publishes | observability.telemetry | TelemetryPublishMessage — one per target per cycle. |
| Consumes | cicd.events | Deployment events, attached to the next snapshot as change markers. |
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/telemetry/latest | The most recent snapshot for each target. |
GET | /api/v1/telemetry/history | Snapshot history (?cluster_id=&from=&to=). |
GET | /api/v1/telemetry/events | Events derived from telemetry. |
GET | /api/v1/telemetry/clusters | Known target identifiers. |
GET | /healthz | Health check. |
Tenant users only see snapshots for their own vClusters.
Configuration
| Variable | Default | Description |
|---|---|---|
COLLECTION_INTERVAL | 30s | How often to collect telemetry. |
TARGET_REFRESH_INTERVAL | 5m | How often to refresh the list of targets. |
K8S_MONITOR_BASE_URL | http://k8s-monitor:8085 | k8s-monitor endpoint. |
SECURITY_API_BASE_URL | http://security-api:8086 | security-api endpoint. |
SECURITY_API_KEY | — | Shared secret for security-api. |
KUBEOPERA_API_BASE_URL | — | Used to discover clusters and CloudSpaces. |
AUTH_SERVICE_BASE_URL | — | Used to obtain the service's own access token. |
SERVICE_CLIENT_ID / SERVICE_CLIENT_SECRET | — | The service identity used for discovery (read-only platform scope). |
AUTH_JWT_ACCESS_SECRET | — | Validates bearer tokens on this service's REST API. |
DATABASE_URL | — | PostgreSQL connection. |
RABBITMQ_URL | — | RabbitMQ connection. |
PORT | 8092 | HTTP port. |
Troubleshooting
- A tenant is missing from
/telemetry/clusters. Check the tenant's CloudSpace isReady; new vClusters appear within one target refresh (five minutes by default). - Snapshots stop arriving. Check the service can reach k8s-monitor and RabbitMQ:
kubectl logs deploy/observability-agent-srv -n kubeopera-core.