Skip to main content
Version: 2.0

SLO Manager

Service: slo-manager · Port: 8100 · Database schema: slo

slo-manager tracks service-level objectives (SLOs): the reliability targets you promise for your services. It evaluates each SLO against live Prometheus data, tells you how much error budget is left and how fast it's burning, and alerts you before you run out.

Key ideas​

TermMeaningExample
SLI (indicator)A measurement between 0 and 1 of good behaviour.Share of requests that succeed.
SLO targetThe minimum SLI you commit to over a rolling window.99.9% over 30 days.
Error budgetHow much badness the target allows: 1 − target.0.1% of requests may fail.
Burn rateHow fast the budget is being used: (1 − SLI) / (1 − target).2.0 = the budget runs out in half the window.

A burn rate of 1 uses the budget exactly over the window. Higher means you'll run out early.

Define an SLO​

In Monitoring → SLOs → New SLO, choose a service and a type, and KubeOpera writes the query for you from your app's standard HTTP metrics. Or use the API:

{
"name": "payments-api availability",
"service": "payments-api",
"namespace": "production",
"type": "availability",
"target": 0.999,
"window_days": 30
}
TypeMeasures
availabilityShare of requests without a 5xx error.
latencyShare of requests faster than a threshold (threshold_ms).
error_rateShare of requests without any error.
throughputWhether request rate stays above a minimum.
customAny PromQL expression you supply in prom_query.

Burn-rate alerts​

slo-manager alerts on multi-window burn rates, the approach recommended by the Google SRE workbook — fast enough to catch sudden outages, calm enough to ignore brief blips:

AlertConditionMeaning
PageBurn rate > 14.4 over both 1h and 5m2% of a 30-day budget gone in an hour.
PageBurn rate > 6 over both 6h and 30m5% gone in six hours.
TicketBurn rate > 1 over both 3d and 6hSteadily trending toward running out.

Alerts go to the incident manager, which opens an incident and notifies your channels.

REST API​

MethodPathDescription
GET · POST/api/v1/slosList or create SLOs.
GET · PUT · DELETE/api/v1/slos/{id}SLO definition.
GET/api/v1/slos/{id}/statusLive SLI, error budget remaining, burn rates and breach state.
GET/api/v1/slos/{id}/historyBudget over time (?from=&to=).
GET/api/v1/slos/summaryTotals: healthy, at risk and breaching.
GET/healthzHealth check.

Configuration​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection.
PROMETHEUS_URLhttp://kube-prometheus-stack-prometheus.monitoring:9090Prometheus.
EVALUATION_INTERVAL1mHow often SLOs are evaluated for alerting.
RABBITMQ_URL—Publishes burn-rate alerts.
AUTH_JWT_ACCESS_SECRET—Validates tokens.
PORT8100HTTP port.