SLO Manager
Service: slo-manager · Port: 8100 · Database schema: slo
slo-manager tracks service-level objectives (SLOs): the reliability targets you promise for your services. It evaluates each SLO against live Prometheus data, tells you how much error budget is left and how fast it's burning, and alerts you before you run out.
Key ideas
| Term | Meaning | Example |
|---|---|---|
| SLI (indicator) | A measurement between 0 and 1 of good behaviour. | Share of requests that succeed. |
| SLO target | The minimum SLI you commit to over a rolling window. | 99.9% over 30 days. |
| Error budget | How much badness the target allows: 1 − target. | 0.1% of requests may fail. |
| Burn rate | How fast the budget is being used: (1 − SLI) / (1 − target). | 2.0 = the budget runs out in half the window. |
A burn rate of 1 uses the budget exactly over the window. Higher means you'll run out early.
Define an SLO
In Monitoring → SLOs → New SLO, choose a service and a type, and KubeOpera writes the query for you from your app's standard HTTP metrics. Or use the API:
{
"name": "payments-api availability",
"service": "payments-api",
"namespace": "production",
"type": "availability",
"target": 0.999,
"window_days": 30
}
| Type | Measures |
|---|---|
availability | Share of requests without a 5xx error. |
latency | Share of requests faster than a threshold (threshold_ms). |
error_rate | Share of requests without any error. |
throughput | Whether request rate stays above a minimum. |
custom | Any PromQL expression you supply in prom_query. |
Burn-rate alerts
slo-manager alerts on multi-window burn rates, the approach recommended by the Google SRE workbook — fast enough to catch sudden outages, calm enough to ignore brief blips:
| Alert | Condition | Meaning |
|---|---|---|
| Page | Burn rate > 14.4 over both 1h and 5m | 2% of a 30-day budget gone in an hour. |
| Page | Burn rate > 6 over both 6h and 30m | 5% gone in six hours. |
| Ticket | Burn rate > 1 over both 3d and 6h | Steadily trending toward running out. |
Alerts go to the incident manager, which opens an incident and notifies your channels.
REST API
| Method | Path | Description |
|---|---|---|
GET · POST | /api/v1/slos | List or create SLOs. |
GET · PUT · DELETE | /api/v1/slos/{id} | SLO definition. |
GET | /api/v1/slos/{id}/status | Live SLI, error budget remaining, burn rates and breach state. |
GET | /api/v1/slos/{id}/history | Budget over time (?from=&to=). |
GET | /api/v1/slos/summary | Totals: healthy, at risk and breaching. |
GET | /healthz | Health check. |
Configuration
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection. |
PROMETHEUS_URL | http://kube-prometheus-stack-prometheus.monitoring:9090 | Prometheus. |
EVALUATION_INTERVAL | 1m | How often SLOs are evaluated for alerting. |
RABBITMQ_URL | — | Publishes burn-rate alerts. |
AUTH_JWT_ACCESS_SECRET | — | Validates tokens. |
PORT | 8100 | HTTP port. |