Learning bite
Recovery and release evidence
Restore a known-good declaration and verify the original data.
On this page
A stopped rollout is not the whole recovery
An abort can stop promotion and restore stable capacity, but it does not automatically reverse a database migration, remove an incompatible queue message, or correct the desired image in Git. Explain which layer each recovery action changes.
For the local exercise, keep the stable image loaded, save the successful transaction case, and record the baseline Git revision before introducing a candidate. Inspect stable and candidate endpoints during recovery rather than trusting a single status label.
Work through three kinds of recovery
For an invented version change A → B, an abort can stop B's progression and return serving capacity to A. A Git change can restore the desired image to A. A database restore would replace data according to a separate recovery procedure. They change different state. If B already wrote a value that A cannot read, returning the image alone is insufficient. This is why the local gate exercise avoids schema changes and fresh writes.
Make a recovery worksheet before the canary lab: known-good image and template, saved case, stable Service selection, current desired Git revision, and the controller currently owning Ledger. After the negative gate, compare each item with its baseline. “Rollout aborted” fills only part of that worksheet.
Small decision exercise: the Rollout is aborted, stable reads pass, and Git still names B. What remains? Record the failed attempt, explicitly choose a corrected desired revision or fix-forward plan, and ensure the next controller handoff will not restore an older unreviewed manifest. If your security lab hardened the container after the first Kubernetes lab, the recovery snapshot must include that hardening. Recovery to an old file can silently undo a later protection.
Keep a release record
| Evidence | Why it matters |
|---|---|
| Desired Git commit and image ID/digest | Identifies what was intended |
| Stable and candidate Pod hashes | Identifies what ran |
| AnalysisRun and Job result | Shows what the gate actually measured |
| Saved case verified through stable | Checks the recovery target |
| Restored desired declaration | Records the accepted target for subsequent release changes |
| Remaining uncertainty | Prevents a lab result becoming a production claim |
An aborted Rollout can remain Degraded while its desired specification still names the candidate. Argo CD does not automatically retry that candidate simply because the same specification is still in Git. For this exercise, explicitly restore the known-good image in Git and reconcile it, or record a deliberate fix-forward decision for a later release. For the independent Rollouts exercise, the module lab pauses Argo CD ownership and restores it only after the declarations agree. Do not let two controllers manage the same workload through conflicting manifests.
Checkpoint
Demonstrate a failing gate, stable recovery, and a passing rerun after restoring the real expected result. Preserve the failed evidence too. If a real assertion fails, investigate its cause rather than loosening the check to obtain a pass. Keep cloud promotion, schema-compatible rollout design, and production approval policy for the platform track.
Checkpoint guide: keep both the failed and successful AnalysisRun records, and verify the original saved case through the restored stable path. The next lab rehearses the full sequence. Its final return to Deployment mode is necessary before the subsequent scale-to-zero fault exercise, which must have only one workload owner.
Sources
Rollouts operations↗, Rollouts and Argo CD FAQ↗, and Argo CD sync controls↗.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.