Learning bite
Observe, recover, and clean up
Finish an experiment with verified recovery and retained evidence.
On this page
Define “back” in terms of the original operation
Recovery has several observations: the desired replica count is restored, a Pod becomes Ready, its Service selects it, and the saved transaction is returned exactly once with the expected balance. These observations build on one another. A fresh transaction passing can show the new path works while leaving the earlier accepted operation unresolved.
Work through an invented record with a failed check at 14:02, replica restoration requested at 14:03, readiness at 14:03:20, and the same case passing at 14:03:30. The observed interval from failed check to verified recovery is 90 seconds. It is not necessarily the full outage duration because the true failure start may precede detection. Keep unknown times unknown.
Practice: list the evidence you would save before cleanup: target and revision, fault and restore times, relevant events/logs, saved case identity, failed result, and same-case recovery result. Separate synthetic identifiers from public summaries. An evidence bundle needs enough context to explain the mechanism, not every log your laptop produced.
Preserve the failing state long enough to understand it
Capture events, workload status, and a limited set of relevant logs before recovery. Use timestamps and a synthetic transaction identifier to connect observations. A screenshot of a red dashboard alone cannot explain the failure path.
Recover the intended workload configuration, wait for readiness, and verify the same saved case. If a transaction was accepted but the test failed before saving its identity, investigate the disposable data rather than claiming recovery from an unrelated fresh transaction.
Clean up controllers in the right order
Stop the local synthetic schedule first and check for active Jobs. Export the useful analysis and incident evidence before deleting their resources. Disable/remove the local Argo CD Application without cascading deletion if you are returning to manually managed resources. Inspect its finalizers first.
Keep storage until you have recorded the outcome. When the entire disposable track is finished, deleting microbank-advanced removes the local cluster and its node-local data. That does not erase Terraform state files on your Mac or clean up a separate Compose project.
Write what the exercise established
Your record should distinguish:
- The injected condition and the observed application behavior.
- Detection latency measured in this run, not a general SLA.
- Recovery action and verification of the original data.
- Remaining gaps, such as missing settlement handling or API authorization.
- The smallest follow-up change and how you will retest it.
The next path can turn these working pieces into platform interfaces, policy enforcement, and cloud adapters. Carry forward the evidence and constraints, including any work still needed before a production deployment.
Cleanup answer: stop producers of new work first, finish/inspect active work, retain the records, then remove the exact disposable resources. Recheck replica count and scheduling after cleanup. Continue to the failure lab with a written recovery criterion; its result should refine the original hypothesis even when the observation differs from your prediction.
Sources
kind cluster deletion↗, Argo CD deletion semantics↗, and postmortem practice↗.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.