Skip to content
← Advanced DevOps

Learning bite

One reversible fault at a time

Choose an injection that reveals a specific dependency.

Documentation reviewed2026-10-01 · 3 min read
On this page

Match the injection to the hypothesis

Choose a fault at the layer you intend to study. Stopping a port-forward breaks the observer's access path. Scaling Ledger to zero interrupts its application process. Blocking only database traffic would test a dependency, but needs an enforcing network setup and a reviewed reversal. These can all make a probe fail for different reasons.

Use this paper comparison: if the question is “does accepted work remain available for Ledger after its consumer returns?”, a brief Ledger interruption is relevant. If you instead delete the entire kind cluster, you also remove local storage, networking, controllers, and the emulator. A failure afterward cannot isolate the queue-consumer question. More severe disruption is not automatically more useful learning.

Before the lab script runs, read its recovery trap and the standalone recovery action: kubectl --context kind-microbank-advanced -n microbank scale deployment ledger --replicas=1. The trap improves ordinary interruption handling, but cannot operate while the host engine or API is unavailable. Keep the command available in another terminal and understand the recorded stop time.

Start with an application-level interruption

Scaling the standalone local Ledger Deployment to zero is enough to study consumer unavailability. With one fault, you can more easily connect the observed symptoms to their cause. Run this only in the lab's documented Deployment mode, after the Rollout resources are removed and the baseline restored.

Do not scale a PostgreSQL Deployment to multiple replicas as a resilience experiment. Shared storage plus multiple independent database processes is not a replication design. Do not delete PVCs to test a normal container restart.

Predict several signals

SignalHypothesis while Ledger is stopped
Ledger Service endpointsNo ready backend
Accounts process checkMay remain successful
New deposit submissionMay be accepted and queued
End-to-end verificationFails or times out
After recoverySaved transaction may become visible; verify it

Measure actual behavior. If Accounts does not accept the transaction, inspect its dependency or configuration and record where the result differs from your prediction.

Escalate only with a reason

Later exercises can test a wrong readiness path, denied network dependency, or bounded latency injection. Each requires a baseline, a tested reversal, and a policy-capable network or appropriate tool. A chaos framework is optional at this stage; the baseline, hypothesis, measurements, and recovery plan matter more than the tool.

Checkpoint: identify the smallest injection that can disprove your hypothesis. Record the affected object, time, observed signals, and reversal. Stop after learning the intended mechanism instead of combining faults until the system becomes impossible to diagnose.

Checkpoint answer: use the scale-to-zero lab only after its Deployment-mode preconditions pass. Compare at least two signals: a process check elsewhere can remain green while Ledger has no ready endpoint. If the injected state immediately disappears, inspect controller ownership before injecting another fault. The next bite explains how to finish with the same case and a useful record.

Sources

Deployment scaling↗, network-policy prerequisites↗, and reliability testing↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.