Observability

SLO Burn-Rate Alerting

ARQQ · September 28, 2026 · 9 min read

Most on-call rotations do not suffer from too little monitoring. They suffer from too many pages that do not require a human. A CPU threshold trips at 3 a.m., the engineer logs in, finds users are unaffected, acknowledges and goes back to sleep. Repeat that a few times a week and the team learns to treat pages as noise, which is exactly the condition under which a real incident gets missed.

Service level objectives (SLOs) and error budgets offer a way out, but only if the alerting is built on them properly. This article walks through a practical path from threshold-based alerting to burn-rate alerting: what to measure, how the maths works, how to wire it into Prometheus and Alertmanager, and where automation should take over from a human.

Why threshold alerts produce fatigue

A static threshold such as "error rate above 1% for five minutes" answers the wrong question. It tells you a number crossed a line; it does not tell you whether that matters against what you promised users. A brief spike on a high-traffic service may be harmless, while a slow, sustained leak of failures can quietly consume a month of reliability without ever tripping the line.

Google's SRE book frames the test every paging rule should pass. Its first question is: "Does this rule detect an otherwise undetected condition that is urgent, actionable, and actively or imminently user-visible?" It follows up by asking whether you can take action in response to the alert, and whether that action is urgent or could be automated (Google SRE Book, Monitoring Distributed Systems). Many cause-based alerts, such as disk, CPU or a single pod restarting, fail that test because they describe a possible cause rather than a user-facing symptom.

The Prometheus project gives the same advice in shorter form: "keep alerting simple, alert on symptoms, have good consoles to allow pinpointing causes" (Prometheus alerting practices). For online serving systems it recommends alerting on "high latency and error rates as high up in the stack as possible."

Step 1: define SLIs that reflect user pain

A service level indicator (SLI) is a ratio of good events to valid events. Start with the four golden signals the SRE book lists (latency, traffic, errors and saturation) but only turn latency and errors into SLIs for paging. Traffic and saturation are better as dashboard context and capacity inputs.

  • Availability SLI: successful requests divided by valid requests, measured at the edge (load balancer or API gateway), not at each microservice.
  • Latency SLI: requests served faster than a chosen threshold divided by valid requests. Use a histogram so the threshold can be changed without re-instrumenting.
  • Freshness or correctness SLI for pipelines and batch jobs: records processed within the expected window divided by records expected.

Measure as close to the user as practical. If you already run distributed tracing, the spans at your ingress tier are a good source; our OpenTelemetry production guide covers rolling that instrumentation out incrementally. Decide explicitly what counts as a "valid" request: health checks, requests rejected for bad input (4xx) and synthetic probes are usually excluded, or they will distort the ratio.

Step 2: turn the SLO into an error budget

An SLO sets a target for the SLI over a window, for example 99.9% of requests succeed over a rolling 30 days. The error budget is simply what is left: 0.1% of valid requests in that window may fail without breaching the objective.

The useful concept for alerting is burn rate, which the Google SRE Workbook defines as "how fast, relative to the SLO, the service consumes the error budget" (SRE Workbook, Alerting on SLOs). A burn rate of 1 means that at the current error rate you will spend exactly the whole budget by the end of the window. For a 99.9% SLO, a sustained 0.1% error rate is a burn rate of 1; a 1.44% error rate is a burn rate of 14.4.

The arithmetic that makes this practical: the fraction of a 30-day budget consumed equals burn rate multiplied by the alert window, divided by 30 days (720 hours).

  • Burn rate 14.4 sustained for 1 hour consumes 14.4 × 1 / 720 = 2% of the budget. Left alone, the full budget is gone in 50 hours.
  • Burn rate 6 sustained for 6 hours consumes 6 × 6 / 720 = 5%. The full budget lasts 5 days.
  • Burn rate 1 sustained for 3 days consumes 72 / 720 = 10%. Not urgent, but it will breach the SLO by the end of the window if nothing changes.

Step 3: use multiwindow, multi-burn-rate alerts

The Workbook evaluates alerting strategies on four attributes: precision (the proportion of alerts that were significant), recall (the proportion of significant events that produced an alert), detection time and reset time, meaning how long an alert keeps firing after the issue is resolved. Single-window alerts force a trade-off between these. Its recommended starting point for a 99.9% SLO balances them with three rules:

  • Page: 2% of budget consumed, long window 1 hour, short window 5 minutes, burn rate 14.4.
  • Page: 5% of budget consumed, long window 6 hours, short window 30 minutes, burn rate 6.
  • Ticket: 10% of budget consumed, long window 3 days, short window 6 hours, burn rate 1.

Each rule fires only when both its long and short windows exceed the burn-rate threshold. The long window provides precision: a single bad minute is not enough. The short window provides fast reset: once the incident is fixed, the short window drops below threshold quickly and the alert clears, instead of lingering for an hour. The Workbook's guidance is to make the short window 1/12 of the long window.

In Prometheus, build this with recording rules so each window's error ratio is computed once and reused. A simplified example for the fast-burn page:

- alert: CheckoutErrorBudgetFastBurn
  expr: |
    (slo:checkout_errors:ratio_rate1h > (14.4 * 0.001))
    and
    (slo:checkout_errors:ratio_rate5m > (14.4 * 0.001))
  labels:
    severity: page
  annotations:
    summary: "Checkout is burning its 30-day error budget at 14x+"
    runbook_url: "https://runbooks.example.internal/checkout/fast-burn"

Two Prometheus details matter here. The for clause makes Prometheus wait a set duration before an alert counts as firing; with a multiwindow rule the short window already provides that damping, so a long for mainly adds detection delay. The newer keep_firing_for clause keeps an alert firing for a period after its condition was last met, which the Prometheus alerting rules documentation notes can prevent flapping and false resolutions. Annotations are the documented place for runbook links, and every paging alert should carry one.

Step 4: route, group and suppress in Alertmanager

Burn-rate rules reduce the number of alert definitions, but a large outage can still produce many notifications. Alertmanager has three documented tools for this (Alertmanager documentation):

  • Grouping "categorizes alerts of similar nature into a single notification." Group by service and SLO, not by pod or instance, so one incident produces one page.
  • Inhibition suppresses notifications for certain alerts while other alerts are already firing. If the fast-burn page is active, inhibit the slow-burn page and the ticket for the same SLO.
  • Silences mute alerts for a set time using matchers. Tie them to change windows and require an expiry, so a forgotten silence does not hide the next real incident.

Route by severity label: page goes to the on-call pager, ticket goes to the team's work queue with no after-hours notification. Cause-based signals such as disk pressure or certificate expiry are still worth tracking, but they belong in tickets or automated remediation, not on the pager, unless they are imminently user-visible.

Step 5: decide what automation handles

The SRE book's question about whether an action "could be automated" is the bridge from alerting to automation. Once pages are tied to user impact, review the remaining cause-based alerts and sort them:

  • Deterministic fix, low blast radius (rotate a full log volume, restart a wedged worker, renew a certificate): automate it and emit an event for the audit trail rather than an alert.
  • Known fix, higher risk (fail over a database, roll back a deploy): automate the diagnosis and prepare the action, but keep a human decision in the loop until the runbook has been exercised enough to trust.
  • Unknown cause: this is what humans are for. The burn-rate page tells them it matters; dashboards and traces help them find why.

We cover the remediation side of this in more depth in Self-Healing Infrastructure. The important sequencing point is that automation should be built on top of a clean, symptom-based alerting layer. Automating responses to noisy alerts only makes the noise faster.

Common pitfalls

Low-traffic services

Burn-rate maths assumes enough events to be statistically meaningful. The Workbook's example: a service receiving 10 requests per hour will show a 10% hourly error rate from a single failed request. Options include synthetic traffic to raise the event count, grouping several small services under one SLO, or lengthening windows and accepting slower detection.

SLOs set to current performance

If the SLO is set to whatever the service happens to achieve today, the error budget stops meaning anything. Set it from what users and dependent teams need, then check whether the system can meet it.

No agreed error budget policy

Alerts tell you the budget is burning; they do not decide what happens next. Agree in advance what the team does when the budget is exhausted, such as pausing risky releases or prioritising reliability work, otherwise every breach becomes a negotiation.

Skipping the review loop

Track precision informally: for each page, record whether it needed human action. Pages that repeatedly did not are candidates for deletion, demotion to a ticket or automation.

A rollout sequence that works

  • Pick one user-facing service with reasonable traffic and an owner willing to try it.
  • Define one availability SLI and one latency SLI at the edge; set 30-day SLOs.
  • Add recording rules and the three burn-rate alerts in parallel with existing alerts, routing them to a test channel for two to four weeks.
  • Compare: which real incidents did each system catch, and which pages were noise?
  • Promote the burn-rate alerts to the pager, demote redundant threshold alerts to tickets, and repeat for the next service.

If you are planning this alongside wider platform work, our automation practice covers the remediation and runbook side.

FAQ

What is an error budget burn rate?

It is how fast a service consumes its error budget relative to the SLO. A burn rate of 1 uses exactly the whole budget over the SLO window; a burn rate of 14.4 on a 30-day window uses 2% of the budget in one hour.

Why use two windows per alert?

The long window avoids paging on brief blips; the short window makes the alert clear quickly once the problem is fixed. The SRE Workbook recommends a short window of 1/12 of the long window.

Should we delete CPU and memory alerts?

Not necessarily. Keep them for capacity planning, dashboards, tickets or automated remediation, but stop paging on them unless they indicate imminent user impact.

Do these numbers work for any SLO?

They are the Workbook's recommended starting point for a 99.9% SLO over 30 days. Recalculate thresholds for your own target and window, then tune based on observed precision and detection time.

← Back to Insights