Skip to content
← Advanced DevOps

Learning bite

Recovery and release evidence

Restore a known-good declaration and verify the original data.

Documentation reviewed2026-10-01 · 3 min read
On this page

A stopped rollout is not the whole recovery

An abort can stop promotion and restore stable capacity, but it does not automatically reverse a database migration, remove an incompatible queue message, or correct the desired image in Git. Explain which layer each recovery action changes.

For the local exercise, keep the stable image loaded, save the successful transaction case, and record the baseline Git revision before introducing a candidate. Inspect stable and candidate endpoints during recovery rather than trusting a single status label.

Work through three kinds of recovery

For an invented version change A → B, an abort can stop B's progression and return serving capacity to A. A Git change can restore the desired image to A. A database restore would replace data according to a separate recovery procedure. They change different state. If B already wrote a value that A cannot read, returning the image alone is insufficient. This is why the local gate exercise avoids schema changes and fresh writes.

Make a recovery worksheet before the canary lab: known-good image and template, saved case, stable Service selection, current desired Git revision, and the controller currently owning Ledger. After the negative gate, compare each item with its baseline. “Rollout aborted” fills only part of that worksheet.

Small decision exercise: the Rollout is aborted, stable reads pass, and Git still names B. What remains? Record the failed attempt, explicitly choose a corrected desired revision or fix-forward plan, and ensure the next controller handoff will not restore an older unreviewed manifest. If your security lab hardened the container after the first Kubernetes lab, the recovery snapshot must include that hardening. Recovery to an old file can silently undo a later protection.

Keep a release record

EvidenceWhy it matters
Desired Git commit and image ID/digestIdentifies what was intended
Stable and candidate Pod hashesIdentifies what ran
AnalysisRun and Job resultShows what the gate actually measured
Saved case verified through stableChecks the recovery target
Restored desired declarationRecords the accepted target for subsequent release changes
Remaining uncertaintyPrevents a lab result becoming a production claim

An aborted Rollout can remain Degraded while its desired specification still names the candidate. Argo CD does not automatically retry that candidate simply because the same specification is still in Git. For this exercise, explicitly restore the known-good image in Git and reconcile it, or record a deliberate fix-forward decision for a later release. For the independent Rollouts exercise, the module lab pauses Argo CD ownership and restores it only after the declarations agree. Do not let two controllers manage the same workload through conflicting manifests.

Checkpoint

Demonstrate a failing gate, stable recovery, and a passing rerun after restoring the real expected result. Preserve the failed evidence too. If a real assertion fails, investigate its cause rather than loosening the check to obtain a pass. Keep cloud promotion, schema-compatible rollout design, and production approval policy for the platform track.

Checkpoint guide: keep both the failed and successful AnalysisRun records, and verify the original saved case through the restored stable path. The next lab rehearses the full sequence. Its final return to Deployment mode is necessary before the subsequent scale-to-zero fault exercise, which must have only one workload owner.

Sources

Rollouts operations↗, Rollouts and Argo CD FAQ↗, and Argo CD sync controls↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.