Skip to content
← Advanced DevOps

Practical lab guide

Lab: interrupt Ledger, recover, and explain the evidence

Run one bounded local fault and verify the same transaction after recovery.

Documentation reviewed2026-10-01 · 4 min read · lab time varies
On this page

Preconditions

Return to the baseline Deployment mode using the end of the canary lab. No Ledger Rollout should remain, and the stable Service must have only the baseline selector. Suspend synthetic scheduling and finish active Jobs. Keep Argo CD reconciliation paused, or remove its Application using non-cascading deletion. Do not run this experiment against kubeadm, EKS, GKE, or another cluster.

Inspect the dedicated namespace annotation, Pods, PVCs, queue state, and image identities. Establish a passing fresh transaction first. Keep Accounts and LocalStack running. Record a maximum disruption window of three minutes as a lab stop condition, not a recovery guarantee.

1. Prepare automatic restoration

Save scripts/study/ledger_fault.sh. It targets one named local Deployment, refuses another starting replica count, records a unique case, and installs a recovery trap before the fault. It expects the Foundations probe and working loopback Accounts port-forward. A Ledger port-forward will terminate when its selected Pod disappears; the post-recovery verification below starts a new one.

bash
#!/usr/bin/env bash
set -euo pipefail
context=kind-microbank-advanced
namespace=microbank
k() { kubectl --context "$context" -n "$namespace" "$@"; }
marker="$(kubectl --context "$context" get namespace "$namespace" \
  -o jsonpath='{.metadata.annotations.learning\.learnwithsk\.dev/local-only}')"
[[ "$marker" == true ]] || { echo 'Missing local study marker' >&2; exit 1; }
replicas="$(k get deployment ledger -o jsonpath='{.spec.replicas}')"
[[ "$replicas" == 1 ]] || { echo 'Restore the single-replica baseline first' >&2; exit 1; }
[[ -z "$(k get rollout ledger --ignore-not-found -o name 2>/dev/null)" ]] || {
  echo 'Remove the Ledger Rollout before this experiment' >&2; exit 1;
}
case_file="evidence/fault-$(date -u +%Y%m%dT%H%M%SZ).json"
mkdir -p evidence
restore() { k scale deployment ledger --replicas=1; }
trap restore EXIT
trap 'exit 130' INT
trap 'exit 143' TERM
k scale deployment ledger --replicas=0
k wait --for=delete pod -l app=ledger --timeout=60s
if python3 scripts/study/probe.py --create --case-file "$case_file"; then
  echo 'Unexpected pass: inspect whether the fault affected the intended path' >&2
  exit 1
fi
[[ -f "$case_file" ]] || {
  echo 'No saved case: inspect Accounts submission before claiming queued recovery' >&2; exit 1;
}
k get pods,events
restore
trap - EXIT
k rollout status deployment ledger --timeout=120s
echo "Restart the Ledger port-forward, then verify: $case_file"

The shell trap cannot recover from every host/process termination, such as SIGKILL or a stopped container engine. Keep the manual recovery command available. The bounded probe returns quickly when Ledger is unreachable; restore the service promptly rather than waiting for a dashboard to update.

2. Run and preserve observations

bash
chmod +x scripts/study/ledger_fault.sh
./scripts/study/ledger_fault.sh

Capture the actual timestamps and output. If the test passes unexpectedly, investigate the target and forwarding process instead of calling the experiment successful. If no case was saved, the exercise did not establish that Accounts accepted and queued a deposit.

Restart the Ledger port-forward in another terminal after the Deployment recovers:

bash
kubectl --context kind-microbank-advanced -n microbank port-forward \
  --address 127.0.0.1 service/ledger 8001:8001

Use the exact case path printed by the script:

bash
: "${FAULT_CASE:?Set the actual evidence/fault-...json path from this run}"
python3 scripts/study/probe.py --verify "$FAULT_CASE"

Expected recovery evidence is the same transaction appearing exactly once with its expected balance. That result must be observed, not assumed. Inspect Accounts outbox, emulator queue, and Ledger logs if it does not converge.

3. Complete the incident record

Record the injected condition, affected component, which health checks stayed green, why the journey check failed, what resumed after restoration, and any remaining uncertainty. Distinguish failure to reach Ledger from a demonstrated queue-processing delay. Include one change that would improve detection and a separate change that would improve recovery.

4. Finish the local track

Verify Ledger is back at one replica, no fault process remains, and synthetic scheduling is suspended. Save useful evidence and redact it before sharing. Remove only the disposable exercise resources when finished; do not prune unrelated Docker images, volumes, or clusters. Destroy the separate emulator resources while the correct instance is reachable before deleting the named kind cluster.

Your completion artifact is a reproducible operating record with actual observations. Carry it into Platform engineering, where advanced GitOps, Kyverno, cloud adapters, and an IDP can build on the verified behavior.

Check what the result supports

If the probe saved a newly accepted operation during the interruption and the same case verifies after Ledger returns, you have evidence for this run's accepted-work recovery. If no case exists, investigate submission before claiming queued recovery. If a different case passes, the original operation remains unresolved.

The recovery trap is a convenience with stated limits, not a substitute for the manual plan. Finish by inspecting the actual replica count, active Jobs, controller owner, and same-case result. Carry the observation and its limits into Platform engineering; repeating the script without those checks would lose the main lesson.

Sources

Kubernetes scaling↗, kubectl wait↗, reliability testing↗, and incident response↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.