Most on-call rotations do not suffer from too little monitoring. They suffer from too many pages that do not require a human. A CPU threshold trips at 3 a.m., the engineer logs in, finds users are unaffected, acknowledges and goes back to sleep. Repeat that a few times a week and the team learns to treat pages as noise, which is exactly the condition under which a real incident gets missed.
Service level objectives (SLOs) and error budgets offer a way out, but only if the alerting is built on them properly. This article walks through a practical path from threshold-based alerting to burn-rate alerting: what to measure, how the maths works, how to wire it into Prometheus and Alertmanager, and where automation should take over from a human.
A static threshold such as "error rate above 1% for five minutes" answers the wrong question. It tells you a number crossed a line; it does not tell you whether that matters against what you promised users. A brief spike on a high-traffic service may be harmless, while a slow, sustained leak of failures can quietly consume a month of reliability without ever tripping the line.
Google's SRE book frames the test every paging rule should pass. Its first question is: "Does this rule detect an otherwise undetected condition that is urgent, actionable, and actively or imminently user-visible?" It follows up by asking whether you can take action in response to the alert, and whether that action is urgent or could be automated (Google SRE Book, Monitoring Distributed Systems). Many cause-based alerts, such as disk, CPU or a single pod restarting, fail that test because they describe a possible cause rather than a user-facing symptom.
The Prometheus project gives the same advice in shorter form: "keep alerting simple, alert on symptoms, have good consoles to allow pinpointing causes" (Prometheus alerting practices). For online serving systems it recommends alerting on "high latency and error rates as high up in the stack as possible."
A service level indicator (SLI) is a ratio of good events to valid events. Start with the four golden signals the SRE book lists (latency, traffic, errors and saturation) but only turn latency and errors into SLIs for paging. Traffic and saturation are better as dashboard context and capacity inputs.
Measure as close to the user as practical. If you already run distributed tracing, the spans at your ingress tier are a good source; our OpenTelemetry production guide covers rolling that instrumentation out incrementally. Decide explicitly what counts as a "valid" request: health checks, requests rejected for bad input (4xx) and synthetic probes are usually excluded, or they will distort the ratio.
An SLO sets a target for the SLI over a window, for example 99.9% of requests succeed over a rolling 30 days. The error budget is simply what is left: 0.1% of valid requests in that window may fail without breaching the objective.
The useful concept for alerting is burn rate, which the Google SRE Workbook defines as "how fast, relative to the SLO, the service consumes the error budget" (SRE Workbook, Alerting on SLOs). A burn rate of 1 means that at the current error rate you will spend exactly the whole budget by the end of the window. For a 99.9% SLO, a sustained 0.1% error rate is a burn rate of 1; a 1.44% error rate is a burn rate of 14.4.
The arithmetic that makes this practical: the fraction of a 30-day budget consumed equals burn rate multiplied by the alert window, divided by 30 days (720 hours).
The Workbook evaluates alerting strategies on four attributes: precision (the proportion of alerts that were significant), recall (the proportion of significant events that produced an alert), detection time and reset time, meaning how long an alert keeps firing after the issue is resolved. Single-window alerts force a trade-off between these. Its recommended starting point for a 99.9% SLO balances them with three rules:
Each rule fires only when both its long and short windows exceed the burn-rate threshold. The long window provides precision: a single bad minute is not enough. The short window provides fast reset: once the incident is fixed, the short window drops below threshold quickly and the alert clears, instead of lingering for an hour. The Workbook's guidance is to make the short window 1/12 of the long window.
In Prometheus, build this with recording rules so each window's error ratio is computed once and reused. A simplified example for the fast-burn page:
- alert: CheckoutErrorBudgetFastBurn
expr: |
(slo:checkout_errors:ratio_rate1h > (14.4 * 0.001))
and
(slo:checkout_errors:ratio_rate5m > (14.4 * 0.001))
labels:
severity: page
annotations:
summary: "Checkout is burning its 30-day error budget at 14x+"
runbook_url: "https://runbooks.example.internal/checkout/fast-burn"
Two Prometheus details matter here. The for clause makes Prometheus wait a set duration before an alert counts as firing; with a multiwindow rule the short window already provides that damping, so a long for mainly adds detection delay. The newer keep_firing_for clause keeps an alert firing for a period after its condition was last met, which the Prometheus alerting rules documentation notes can prevent flapping and false resolutions. Annotations are the documented place for runbook links, and every paging alert should carry one.
Burn-rate rules reduce the number of alert definitions, but a large outage can still produce many notifications. Alertmanager has three documented tools for this (Alertmanager documentation):
Route by severity label: page goes to the on-call pager, ticket goes to the team's work queue with no after-hours notification. Cause-based signals such as disk pressure or certificate expiry are still worth tracking, but they belong in tickets or automated remediation, not on the pager, unless they are imminently user-visible.
The SRE book's question about whether an action "could be automated" is the bridge from alerting to automation. Once pages are tied to user impact, review the remaining cause-based alerts and sort them:
We cover the remediation side of this in more depth in Self-Healing Infrastructure. The important sequencing point is that automation should be built on top of a clean, symptom-based alerting layer. Automating responses to noisy alerts only makes the noise faster.
Burn-rate maths assumes enough events to be statistically meaningful. The Workbook's example: a service receiving 10 requests per hour will show a 10% hourly error rate from a single failed request. Options include synthetic traffic to raise the event count, grouping several small services under one SLO, or lengthening windows and accepting slower detection.
If the SLO is set to whatever the service happens to achieve today, the error budget stops meaning anything. Set it from what users and dependent teams need, then check whether the system can meet it.
Alerts tell you the budget is burning; they do not decide what happens next. Agree in advance what the team does when the budget is exhausted, such as pausing risky releases or prioritising reliability work, otherwise every breach becomes a negotiation.
Track precision informally: for each page, record whether it needed human action. Pages that repeatedly did not are candidates for deletion, demotion to a ticket or automation.
If you are planning this alongside wider platform work, our automation practice covers the remediation and runbook side.
It is how fast a service consumes its error budget relative to the SLO. A burn rate of 1 uses exactly the whole budget over the SLO window; a burn rate of 14.4 on a 30-day window uses 2% of the budget in one hour.
The long window avoids paging on brief blips; the short window makes the alert clear quickly once the problem is fixed. The SRE Workbook recommends a short window of 1/12 of the long window.
Not necessarily. Keep them for capacity planning, dashboards, tickets or automated remediation, but stop paging on them unless they indicate imminent user impact.
They are the Workbook's recommended starting point for a 99.9% SLO over 30 days. Recalculate thresholds for your own target and window, then tune based on observed precision and detection time.