Skip to content
← Platform engineering

Learning bite

Controller access and recovery

Plan recovery for the delivery system as well as for the application it manages.

Documentation reviewed2026-10-01 · 3 min read
On this page

The delivery system has dependencies too

A controller needs access to its source, the Kubernetes API and the credentials required for its work. Its failure can stop new changes while already-running application processes continue. Diagnose the operation that failed before deciding whether users are affected.

Controllers become a dependency

If Argo CD cannot reach Git, a healthy running workload may continue serving while new changes cannot reconcile. If a policy webhook is unavailable, new admission requests may fail depending on its configuration. Treat those as platform incidents with different symptoms from an application outage.

Record the installed release, configuration source, credentials owner, and recovery procedure for each controller. Pin releases when repeating an exercise. An unversioned install URL is not a reproducible platform definition.

Recover the right state

Argo CD configuration backup is different from a PostgreSQL backup. Restoring Applications can reconnect reconciliation, but it does not restore application data or prove repository credentials remain valid. Exported controller configuration may include sensitive material; store it privately and test restoration in a separate disposable environment.

For this local exercise, capture declarative project/application files in Git and keep private credential recovery instructions elsewhere. A single-node kind cluster remains disposable and non-HA.

Use a recovery decision table

A diagnostic map separates source fetch, render, authorization, admission and workload failures.

Open diagram at full size.

Start at the first failing stage rather than restarting everything. If Argo cannot fetch Git, check the source URL, revision, network path and credential validity. If it fetches but cannot render, inspect the overlay path and referenced files. If the API denies the write, distinguish RBAC denial from an admission-policy rejection. If the objects are accepted but unhealthy, use the workload's events, logs and application checks.

For a read-only local observation, after the GitOps lab exists, run:

bash
kubectl --context kind-microbank-advanced -n argocd get applications
kubectl --context kind-microbank-advanced -n microbank get deployments,pods

These commands show different layers. An Application condition is not a database health check. Record the installed controller release and the revision it was trying to use before planning a recovery.

Rehearse restoring the definition of an Application from Git without actually deleting the controller. Then list what that definition does not restore: credentials, application database contents and any uncommitted operational knowledge. That list becomes the backup/recovery plan, with private material kept outside the public course.

Try it

Without stopping the controller, rehearse the decision process for repository unavailability: inspect Application conditions, compare the last observed revision with the intended revision, check repository connectivity, and determine whether workloads are affected. Then write the smallest recovery sequence and the checks that would confirm it worked.

Inspect the permissions of the controller identity through Kubernetes RBAC. Distinguish permissions required for the supported resource kinds from broad installation defaults. Never remove permissions from a shared controller just to demonstrate a lesson.

Checkpoint and revision

Explain what remains usable if the portal, Git host, Argo CD, or policy engine is unavailable. Include who can perform an emergency action, how it is recorded, and how desired state is reconciled afterward. Include the observations that confirm restoration, even when the recovery action is a restart.

Compare your reasoning

If the portal is unavailable, the reviewed Git workflow may remain usable. If Git is unavailable, already-running Pods may remain usable. Neither observation guarantees future recovery. After an emergency manual change, reconcile the intended state before normal automation resumes.

Sources

Argo CD disaster recovery↗, Kyverno availability↗, Kubernetes RBAC↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.