Skip to content
← Advanced DevOps

Practical lab guide

Lab: define an SLI and write a recovery record

Calculate an explicit fixture, then replace it with measured local evidence.

Documentation reviewed2026-10-01 · 3 min read · lab time varies
On this page

1. Define the measurement contract

Write a one-page SLI definition for fresh synthetic deposit journeys. Include creation/acceptance, idempotent repetition, Ledger verification, timeout budget, eligible events, and observation window. Record that it is a local API journey with synthetic data, not a production customer SLO.

Do not mix periodic reads of an existing case into the fresh-deposit denominator. Keep missing checks and host-sleep periods visible. Record source revision and competing host workloads with your measurement.

2. Check arithmetic with a labelled fixture

Create evidence/sli-fixture.json with this invented test dataset, used only to test the calculation:

json
{"eligible": 1000, "failed": 7, "missing": 3, "objective": 0.99}

Save scripts/study/sli_report.py:

python
import json
import math
import sys
from pathlib import Path


def report(data):
    eligible, failed, missing = (data[key] for key in ('eligible', 'failed', 'missing'))
    target = data['objective']
    if any(type(value) is not int or value < 0 for value in (eligible, failed, missing)):
        raise ValueError('Event counts must be nonnegative integers')
    if eligible == 0 or failed > eligible:
        raise ValueError('A nonempty, consistent observed population is required')
    if isinstance(target, bool) or not isinstance(target, (int, float)):
        raise ValueError('Objective must be numeric')
    if not math.isfinite(target) or not 0 < target < 1:
        raise ValueError('Objective must be finite and between zero and one')
    allowance = eligible * (1 - target)
    return {'observed_success_fraction': (eligible - failed) / eligible,
            'allowed_failures': allowance, 'remaining_budget': allowance - failed,
            'burn_rate': (failed / eligible) / (1 - target),
            'missing_observations': missing, 'evidence_complete': missing == 0}


if __name__ == '__main__':
    print(json.dumps(report(json.loads(Path(sys.argv[1]).read_text())), indent=2))
bash
python3 scripts/study/sli_report.py evidence/sli-fixture.json

Expected fixture values: observed success 0.993, allowed failures approximately 10, remaining approximately 3, burn rate approximately 0.7, and evidence_complete=false. Floating-point output may contain tiny rounding differences. Missing observations remain a separate gap; this script does not decide whether your eventual SLO policy counts them as failures.

Test zero observations, failures exceeding eligible events, negative counts, and a non-finite objective. Each test should reject the invalid input.

3. Collect a small real sample

Run a bounded set of local fresh journeys sequentially, using a unique case file for every run. Record start/end times and failures, including partially completed cases. Stop if the host is under pressure. Replace the fixture with your measured counts and explain the small sample's limitations. The fixture’s 99.3% is a calculation example; report your application’s measured result separately.

4. Practise an incident record

Use the later controlled-failure lab, or the observer-access interruption from the observability lab. Capture baseline, detection, diagnosis, recovery, and verification. Distinguish the affected component precisely. An interrupted port-forward is not a Ledger outage.

Checkpoint

Publish the SLI definition, actual observations, calculation, incident timeline, and one follow-up with a verification method. Keep sensitive/raw case evidence private. Read the record a week later: could another learner reproduce the experiment without guessing which image or environment you used?

Check the denominator before reporting

In the fixture, eligible is the observed population and missing is a separate collection gap. The report therefore computes observed success, not a policy-adjusted score for all planned checks. Decide how missing runs belong in your actual SLI definition before interpreting a real result. Do not mix the read-only exporter runs into a fresh-deposit population.

For the incident record, pair every conclusion with its supporting observation. A port-forward interruption demonstrates observer failure; a later controlled Ledger stop demonstrates a different mechanism. Both are valid lab records when their scope is explicit.

Sources

Implementing SLOs↗, error-budget policy↗, and postmortem culture↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.