Skip to content
← Advanced DevOps

Learning bite

Error budgets and alerting

Connect measured failure consumption to an explicit release decision.

Documentation reviewed2026-10-01 · 3 min read
On this page

Calculate consumption, then choose an action

For a count-based objective, the error budget is the fraction of eligible events allowed to fail. Burn rate compares the observed failure fraction with that allowance. If a hypothetical objective is 99% and 5% of eligible events fail in a window, the burn rate is 5. That describes the window; it does not tell you the root cause.

python
objective = 0.99
observed_failure_fraction = 0.05
burn_rate = observed_failure_fraction / (1 - objective)
print(round(burn_rate, 2))

The fixture prints 5.0. Do not use this example's target or alert threshold as a production recommendation. Small local samples can move sharply after one failure.

Interpret the rate and the amount together

At a hypothetical 99% objective, the allowed bad fraction is 1%. Seven failures in 1,000 observed events consume 7 of the allowed 10 failures, leaving 3; the observed bad fraction is 0.7%, so burn rate over that same population is 0.7. Fifty failures in 1,000 would produce a burn rate of 5 and exceed the budget. A burn rate of 5 means five times the allowed bad fraction during the measured window, not five failures and not necessarily five minutes until exhaustion.

Now imagine a long window that still looks acceptable but a recent short window that is failing quickly. A multi-window alert asks both whether consumption is material over a longer period and whether the problem is still active in a shorter one. A shorter window can detect change sooner but is noisy with little traffic. A longer window smooths noise but can retain an old incident after recovery. Thresholds must connect to a chosen budget, window, and response urgency.

Guided decision: use a worksheet with 1,000 observed events, seven failures, three missing observations, and a 99% training objective. Calculate 99.3% observed success, 10 allowed failures, 3 remaining, and burn rate 0.7. Do the three missing checks prove success? No; they are incomplete evidence until the measurement policy resolves them. If a candidate check is presently failing, would the remaining long-window budget require you to promote it? No. A release gate asks a different, immediate question.

Separate release analysis from long-window reliability

A short canary check decides whether to continue one rollout. An SLO evaluates reliability over a defined population and window. A passing canary does not replenish a spent budget; a long-window SLO can hide a fresh, severe regression.

Production alert designs often combine windows to detect fast and sustained budget consumption. In this course, first prove the single-window arithmetic and missing-data behavior. Record a policy such as “pause the lab release and investigate after a failed candidate journey” and label it as a training policy rather than a statistically validated SLO policy.

Checkpoint

For the SRE lab fixture, calculate observed success, allowed failures, remaining budget, and burn rate. State which failures are application failures and which are missing evidence. Then propose a response: halt a new release, recover the known-good version, or improve the measurement. An error budget is a decision tool, not permission to ignore a user-visible incident.

Response guide: choose an action that follows from the observation. A failed candidate can pause that release; an ongoing user failure needs mitigation; an unavailable observer needs measurement repair and explicit uncertainty. A low burn rate cannot prove recovery while the observation path is broken. Carry that distinction into the incident timeline in the next bite.

Sources

Alerting on SLOs↗, example error-budget policy↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.