Skip to content
← Advanced DevOps

Learning bite

Baseline and blast radius

Turn a fault idea into a bounded local experiment.

Documentation reviewed2026-10-01 · 3 min read
On this page

Choose what you want to learn from a failure

A controlled failure experiment changes one condition to test a prediction. The baseline describes behavior immediately before the change. The blast radius is the set of resources and users that can be affected. The stop condition tells you when to end early. A useful experiment is small enough that its observations can support or contradict the prediction.

For a paper example, stopping Ledger removes the component that consumes queued work and serves Ledger reads. Accounts may still accept a deposit because acceptance and Ledger completion occur in different processes. That expectation depends on the actual application and queue setup, so first identify what durable record would demonstrate acceptance. A request error before acceptance tests a different part of the path.

Write a six-line experiment card before the live lab: hypothesis; target context/namespace/Deployment; baseline check; maximum duration; manual recovery; same-case verification. For this track the maximum planned disruption is three minutes, the manual action restores Ledger to one replica, and verification checks the transaction created during the interruption. If that transaction was never saved, do not silently substitute another one.

Write a falsifiable hypothesis

For MicroBank: “If Ledger processing stops briefly, Accounts may still accept a deposit, the end-to-end check will fail to converge, and processing should resume after Ledger returns.” This is a hypothesis to test, not a promised outcome.

The baseline includes a passing fresh transaction, the actual image revisions, resource usage, queue configuration, and a saved case. Without a recorded baseline, you could mistake an existing problem for an effect of the injected fault.

Define the boundary before acting

Use only the disposable microbank namespace in kind-microbank-advanced. Use synthetic data. Specify the single target, maximum duration, stop condition, recovery command, and verification. Do not interrupt your Mac's network, kill the container engine, or delete the cluster as a first experiment.

Check which controller owns the target. Argo CD self-healing or a Rollout can undo an injected scale change. Pause reconciliation deliberately or choose a lab state with a single Deployment owner. A controller repairing your fault is an observation, not evidence that the fault was never injected.

Preflight questions

  • Is the baseline passing immediately before the experiment?
  • Can I name the exact resource and context affected?
  • Do I have a recovery path that does not require the failed component?
  • What observation will cause an early stop?
  • How will I verify the original transaction after recovery?

Checkpoint: write the experiment card before running anything. Keep the failure small and reversible so you can connect each observation to the mechanism you are testing.

Preflight answer: a broken baseline invalidates comparison because the fault's effect cannot be isolated. An active controller that immediately restores replicas also changes the experiment. Resolve both before injection. Continue by choosing the smallest reversible fault that can test your card.

Sources

Google SRE testing for reliability↗, Principles of Chaos Engineering↗, and MicroBank recovery exercise.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.