Learning bite
SLIs and justified SLOs
Define a measurable user outcome and state the limits of local evidence.
On this page
Turn “reliable” into a measurement
A service level indicator (SLI) measures a defined aspect of service. A service level objective (SLO) sets a target for that indicator over a stated population and window. A service level agreement (SLA) is a commitment with agreed consequences; this local exercise creates no SLA. Start with what a caller needs, then decide how to measure it.
Suppose a deposit API accepts 100 requests and returns quickly, but only 92 produce the expected Ledger result within the chosen deadline. An acceptance-only indicator can look perfect while the complete-journey indicator is 92%. That difference is useful: it tells you the indicator must state whether it measures acceptance, eventual result, or both. The deadline is part of the success definition, not a value added after seeing the results.
Practice before calculating: write an event definition for ten invented scheduled checks: eight pass, one fails, one never runs. If the policy measures completed checks, observed success is 8/9 with one missing observation reported separately. If the policy promises scheduled successful runs and treats a missed run as bad, it is 8/10. Neither denominator can be selected after the fact merely to improve the score. Keep the collection gap visible whichever policy you choose.
Define success before calculating it
For this lab, a successful fresh-deposit journey creates a synthetic account, accepts a deposit, returns the same transaction ID for a repeated idempotency key, and observes exactly one corresponding Ledger entry with the expected balance. A successful HTTP /health request is a different indicator.
Write the numerator, denominator, measurement point, evaluation window, latency condition, and exclusions. Decide how timeouts, missing runs, host sleep, and telemetry failure are represented. Keep failures in the eligible population so the indicator reflects what users experience.
Use a learning target honestly
A hypothetical 99% success objective over 1,000 eligible local journeys allows 10 failures. This is arithmetic practice, not a justified production target. A short, periodic local probe samples one controlled workflow, not every customer, location, or real traffic mix.
eligible = 1000
failed = 7
target = 0.99
observed_success = (eligible - failed) / eligible
allowed_failures = eligible * (1 - target)
remaining = allowed_failures - failed
print(round(observed_success, 4), round(remaining, 2))
Expected fixture values are 0.993 and approximately 3.0. Count-based and time-based objectives are not interchangeable without an explicit model. Treat zero observations as missing data, not as a successful measurement.
Apply it to MicroBank
Record API acceptance separately from eventual Ledger convergence. The source does not implement the Accounts settlement consumer, so status=settled is not a valid success criterion for this baseline. Browser UI mocks are also not transaction evidence.
Checkpoint: write one defensible lab SLI, identify at least two blind spots, and explain what production evidence and stakeholder decisions would be needed to choose a real objective.
Checkpoint guide: a defensible local definition states “fresh journeys begun during this recorded window that meet these completion checks within this deadline,” counts failures, and reports collection gaps separately. Its blind spots include the local host/network and the limited synthetic workload. Historical performance, user expectations, consequence of failure, and achievable engineering trade-offs are needed before choosing a production target. Next, translate that target into an error budget.
Sources
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.