Skip to content
← DevOps foundations

Practical lab guide

MicroBank 7: break, recover, and close the lab

Verify a stopped consumer, preserved database data, and scoped infrastructure cleanup.

Documentation reviewed2026-10-01 · 5 min read · lab time varies
On this page

Start with a known successful case

Complete steps 1–6 and retain evidence/first-case.json. The project milestone is the implemented local API flow, not a production banking system. Keep a short runbook with source SHA, image IDs, LocalStack/provider versions, startup sequence, probe command, shutdown procedure, and known application gaps.

Break only your Ledger consumer

Stop the Ledger service in this named Compose project, then submit a new synthetic case:

bash
./scripts/study/compose.sh stop ledger
python3 scripts/study/probe.py --create --case-file evidence/recovery-case.json
./scripts/study/compose.sh logs --tail=60 accounts
./scripts/study/compose.sh start ledger
./scripts/study/compose.sh ps
python3 scripts/study/probe.py --verify evidence/recovery-case.json

The create operation should save its case then fail verification while Ledger is unavailable. If case creation itself failed, inspect Accounts before continuing; do not assume a file was written. After Ledger starts and becomes responsive, rerun the read-only verification until the bounded check succeeds or investigate the error. Trace the outbox and queue while it is stopped. Do not receive/delete messages from the live consumer queue just to inspect them: that changes visibility and can interfere with the exercise.

This is a deliberate service interruption, not an automated chaos framework. A healthy process alone does not prove the queued transaction reached the Ledger.

Interpret the interruption without creating extra work

When Ledger is stopped, Accounts may still accept and publish a new request. A failed verification then means the complete result was not observed within the probe's bounds; it does not prove Accounts rejected the write. This is why you retain the created case and verify it after restarting the consumer.

The expected recovery is that Ledger processes the queued event and the saved case's checks pass. If that does not happen, inspect one handoff at a time: Accounts outbox, topic/subscription, queue visibility and policy, consumer logs, then Ledger data. Preserve the case identifier so every observation refers to the same operation. Reading messages through a competing consumer can change visibility, so prefer appropriate queue attributes and logs during diagnosis.

Checkpoint: why not rerun --create to check recovery? It creates another account and deposit rather than checking the interrupted case. Why can a health endpoint pass while verification fails? Process health does not establish consumption or the expected account balance.

Verify persistence before tearing down

bash
./scripts/study/compose.sh restart accounts ledger
./scripts/study/compose.sh ps
python3 scripts/study/probe.py --verify evidence/first-case.json
python3 scripts/study/probe.py --verify evidence/recovery-case.json

Wait for service readiness and repeat the verification after transient startup failures. Both cases should retain their independent account balances and matching entries. Named volumes preserve database files, while SPRING_JPA_HIBERNATE_DDL_AUTO=update avoids the source's local create-drop behavior. This is restart evidence, not a tested backup/restore or a production data-durability guarantee.

For a normal pause, use ./scripts/study/compose.sh stop. On resume, start LocalStack/databases first, verify or reconcile Terraform, regenerate the Ansible configuration, and then start the services. If emulator persistence is unavailable in your chosen edition/version, document that limitation: resource re-creation does not recover previously queued messages.

Review cleanup in reverse ownership order

Stop application publishers/consumers before destroying their dependencies. Keep LocalStack running until Terraform finishes:

bash
./scripts/study/compose.sh --profile ui stop frontend accounts ledger
terraform -chdir=infra/study/terraform plan -destroy -out=cleanup.tfplan
terraform -chdir=infra/study/terraform show cleanup.tfplan
terraform -chdir=infra/study/terraform apply cleanup.tfplan
terraform -chdir=infra/study/terraform state list
./scripts/study/compose.sh --profile ui down

Inspect the destruction plan before applying it. It should contain only this lab's resources. If you added objects to the practice bucket, remove only the disposable objects you own before destroying it. state list should be empty after successful destruction. Retain the state/lock files until this has been verified.

The normal down command retains the named database and emulator volumes. When deliberately discarding all synthetic lab data, run ./scripts/study/compose.sh --profile ui down --volumes for this project. This erases its named volumes; it is not part of a persistence or rollback check. Do not use broad Docker pruning. Remove only the lab image tags you no longer need, and keep source and sanitized evidence.

Publish the learning evidence

Add a page to the personal learning site you built in System engineer foundations:

EvidenceWhat it demonstrates
Versioned Compose profile and image IDsExplicit packaging, networking, and runtime inputs
Reviewed Terraform plan and no-change follow-upResource ownership and change review
Ansible second runConfiguration convergence
Probe tests and actual local resultBounded automation and the implemented API contract
Consumer interruption and recovery recordDiagnosis through asynchronous dependencies
Restart verification and cleanup recordDeliberate lifecycle and data handling
CI run and deployment noteTraceability from source to selected artifact

Keep a visible backlog: backend authorization, input validation, concurrent idempotency constraints, settlement consumption, real notification service, frontend balance integration, database migrations, and stronger automated tests. Record observations instead of marking these features complete. Synthetic personal information, tokens, raw environment files, and state do not belong in the published evidence.

Continue to Advanced DevOps for Kubernetes, GitOps, observability, progressive delivery and broader resilience exercises. Real AWS/EKS/GKE adapters need a separate identity, networking, security and cost design. The foundations implementation stays local.

Sources and practice status

Compose lifecycle commands↗, Terraform destroy planning↗, MicroBank application source↗.

These guides supply implementation files and acceptance steps. Documentation review, static checks, unit tests, and a site build are not substitutes for running the full stack with the chosen LocalStack image/authentication and recording the results. Keep reviewed status until that execution evidence is available.

Author verification record

On 30 September 2026, the seven probe unit tests and CLI failure cases passed. The extracted Ansible playbook rendered disposable local fixture files, preserved mode 0600, reported changed=0 on its second run, and rejected a non-local queue contract. Compose configuration parsing confirmed the intended six services and loopback publications. Terraform 1.12.2 validated the configuration with the signed AWS provider 6.66.0. Python/Bash snippets passed syntax checks.

These checks did not start the MicroBank containers, apply resources to LocalStack, run hosted CI, or verify Auth0 login. Run the documented acceptance steps with your selected image/authentication setup before adding end-to-end execution evidence.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.