Dashboards Reference
The dashboard's Monitoring section has a page for each kind of question. This reference helps you pick the right one.
| Page | Answers | Data from |
|---|---|---|
| Performance | Is my infrastructure healthy? CPU, memory, network and pod status. | The configured metrics source (METRICS_SOURCE). |
| APM | Is this service slow or failing? Request rate, latency and error rate per service and endpoint. | apm-gateway (Prometheus). |
| Traces | Where is time spent in this request? End-to-end traces and the service map. | tracing-gateway (Jaeger). |
| SLOs | Are we meeting our objectives? Compliance, error budget and burn rate. | slo-manager. |
| Cost | What does this cost? Spend by cluster, namespace and workload, with forecasts. | k8s-monitor. |
| Security | Where are we exposed? Posture score, findings and vulnerabilities. | security-api. |
| Optimizer | What's over- or under-provisioned? Right-sizing suggestions per workload. | k8s-monitor. |
| AI Optimizer | What does continuous AI analysis recommend? Health assessment and optimization reports. | kubeopera-ai. |
| Actions | What has automation done? Every automated remediation, its reason and its measured outcome. | action-agent and feedback-agent. |
Choosing between similar pages
- Performance vs APM — "is my node under memory pressure?" is Performance; "is my checkout service slow?" is APM.
- Optimizer vs AI Optimizer — Optimizer lists concrete right-sizing changes per workload; AI Optimizer gives a broader, AI-written assessment of the cluster and suggestions across areas.
- APM vs Traces — APM tells you that an endpoint is slow; Traces tells you where inside the request the time goes.
Keeping data fresh
Every page updates on its own every 30–60 seconds, and live views (agent runs, logs) stream in real time. Use the time-range and cluster selectors at the top of each page to change what you're looking at.
Next steps
- Monitoring — instrument your apps so these pages have data.
- Advanced monitoring — the same data from scripts and AI tools.