Learning bite
One reversible fault at a time
Choose an injection that reveals a specific dependency.
On this page
Match the injection to the hypothesis
Choose a fault at the layer you intend to study. Stopping a port-forward breaks the observer's access path. Scaling Ledger to zero interrupts its application process. Blocking only database traffic would test a dependency, but needs an enforcing network setup and a reviewed reversal. These can all make a probe fail for different reasons.
Use this paper comparison: if the question is “does accepted work remain available for Ledger after its consumer returns?”, a brief Ledger interruption is relevant. If you instead delete the entire kind cluster, you also remove local storage, networking, controllers, and the emulator. A failure afterward cannot isolate the queue-consumer question. More severe disruption is not automatically more useful learning.
Before the lab script runs, read its recovery trap and the standalone recovery action: kubectl --context kind-microbank-advanced -n microbank scale deployment ledger --replicas=1. The trap improves ordinary interruption handling, but cannot operate while the host engine or API is unavailable. Keep the command available in another terminal and understand the recorded stop time.
Start with an application-level interruption
Scaling the standalone local Ledger Deployment to zero is enough to study consumer unavailability. With one fault, you can more easily connect the observed symptoms to their cause. Run this only in the lab's documented Deployment mode, after the Rollout resources are removed and the baseline restored.
Do not scale a PostgreSQL Deployment to multiple replicas as a resilience experiment. Shared storage plus multiple independent database processes is not a replication design. Do not delete PVCs to test a normal container restart.
Predict several signals
| Signal | Hypothesis while Ledger is stopped |
|---|---|
| Ledger Service endpoints | No ready backend |
| Accounts process check | May remain successful |
| New deposit submission | May be accepted and queued |
| End-to-end verification | Fails or times out |
| After recovery | Saved transaction may become visible; verify it |
Measure actual behavior. If Accounts does not accept the transaction, inspect its dependency or configuration and record where the result differs from your prediction.
Escalate only with a reason
Later exercises can test a wrong readiness path, denied network dependency, or bounded latency injection. Each requires a baseline, a tested reversal, and a policy-capable network or appropriate tool. A chaos framework is optional at this stage; the baseline, hypothesis, measurements, and recovery plan matter more than the tool.
Checkpoint: identify the smallest injection that can disprove your hypothesis. Record the affected object, time, observed signals, and reversal. Stop after learning the intended mechanism instead of combining faults until the system becomes impossible to diagnose.
Checkpoint answer: use the scale-to-zero lab only after its Deployment-mode preconditions pass. Compare at least two signals: a process check elsewhere can remain green while Ledger has no ready endpoint. If the injected state immediately disappears, inspect controller ownership before injecting another fault. The next bite explains how to finish with the same case and a useful record.
Sources
Deployment scaling↗, network-policy prerequisites↗, and reliability testing↗.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.