Dashboard
/dashboard is the platform-admin landing page after login — the tenant-customer equivalent is /home, a different page covered in Your First Login. It aggregates real signals from k8s-monitor and Prometheus into four sections, refreshed together on one timer rather than each having its own.
Cluster Overview
Six metric cards, all sourced from a single GET /api/k8s-monitor/combined call (fetchK8sMonitorReport('combined')):
| Card | What it shows |
|---|---|
| Health Score | monitorReport.health.healthScore out of 100, plus a count of active issues |
| Ready Nodes | readyNodes/totalNodes from the same report |
| Restarting Pods | Pod restart count, with running/total pod counts as a subtitle |
| CPU Usage | Average CPU utilization across nodes (from the data-source layer — see below) |
| Memory Usage | Current memory usage percentage, or a fallback average when no live snapshot is available |
| Monthly Cost | Computed client-side as sum(node.TotalCost for every node) × 24 × 30 — an hourly-rate extrapolation, not a billed monthly total |
Active Health Issues
Shown only when monitorReport.health.issues is non-empty. Each card shows a severity badge (critical/warning/info), the affected resource, a message, and — when the backend supplies one — a suggested fix. Limited to the first 6 issues; there's no "view all" beyond that today.
Resource Utilization
CPU and memory history charts, both sourced from k8s-monitor's historical data. When k8s-monitor returns no historical points for either metric, the panel shows an explicit "Not Available" state rather than an empty or fabricated chart.
Observability
Four more panels — Latency, Pod Status, Network Traffic, and Error Count — sourced from Prometheus via lib/observabilityApi.ts (Pod Status instead comes from the same metrics data source as the Cluster Overview cards, driven by METRICS_SOURCE). Each degrades honestly rather than fabricating data:
- If Prometheus itself isn't reachable, the panel explains that workloads need to be instrumented with a Prometheus client library.
- If Prometheus is reachable (proven by Latency/Network already having real data) but the Error panel comes back empty, the panel explicitly distinguishes "no 5xx errors were recorded" from "nothing is instrumented" — PromQL can't tell those apart on its own from an empty query result, so the frontend makes the distinction using the other panels' own reachability as evidence.
Error-rate coverage today is limited to the one or two backend services actually instrumented with HTTP metrics — not a platform-wide signal yet.
Refresh behavior
All three data sources (metrics, the k8s-monitor combined report, and observability) are fetched together, on mount and then on a single shared 60-second setInterval — not per-section independent timers. There's no stale-data indicator distinct from the "Updated" / "Last refresh" timestamps already shown in the Cluster Overview header.
What this page doesn't have (yet)
An earlier version of this page described a Cost Trend chart, a Recent Anomalies feed, a CI/CD Activity feed, and an AI Agent Activity feed — none of these exist on /dashboard today. Recent anomalies, pipeline activity, and agent runs are real, but live on their own dedicated pages (Analytics, Pipelines, AI Agents) rather than being summarized here.