Skip to content
← Advanced DevOps

Practical lab guide

Lab: gate a local Ledger canary with synthetic checks

Run scheduled local reads and make candidate-specific analysis block a bad promotion.

Documentation reviewed2026-10-01 · 9 min read · lab time varies
On this page

Scope and prerequisites

The Kubernetes baseline must pass a fresh transaction and preserve evidence/kind-case.json. Stop the earlier probe exporter to reduce noise. Keep Accounts at one replica. This lab checks Ledger's read behavior against a saved case; it does not prove new writes, browser login, database migrations, or concurrent consumer correctness.

Stable and candidate Ledger processes share the disposable database and queue. Do not create new transactions during this exercise. Verify the queue has drained first. A production design needs explicit schema compatibility and consumer-concurrency tests before sharing those dependencies across versions.

Use only kind-microbank-advanced. These manifests live in a local exercise directory, never a cloud overlay or hosted CI schedule. Tool versions and local memory usage belong in the evidence record. With two desired Ledger replicas and one allowed surge, up to three Ledger Pods may run temporarily; stop optional stacks first and do not continue under host memory pressure.

1. Hand off workload ownership

If you completed GitOps, inspect the microbank-local Application's finalizers and remove it using the documented non-cascading procedure from the previous lab. Its Deployments should remain. Before handoff, compare the latest GitOps base, the original generated manifest, and your hardening record. Save one reviewed Kubernetes List containing the two API Deployments and Services as .local/canary-baseline.json. Preserve the current image and security changes. Use this reviewed snapshot as both the input and recovery source for the exercise so an older copy cannot undo your changes. Do not leave a GitOps Application recreating Ledger's Deployment while a Rollout manages Ledger.

Install a reviewed Argo Rollouts release and matching CLI plugin using the official installation guide. Download the exact release manifest and inspect it before applying:

bash
: "${ROLLOUTS_VERSION:?Select an explicit reviewed Argo Rollouts release tag}"
curl --fail --location --output .local/argo-rollouts.yaml \
  "https://github.com/argoproj/argo-rollouts/releases/download/$ROLLOUTS_VERSION/install.yaml"
shasum -a 256 .local/argo-rollouts.yaml
kubectl --context kind-microbank-advanced create namespace argo-rollouts
kubectl --context kind-microbank-advanced apply -n argo-rollouts --server-side -f .local/argo-rollouts.yaml
kubectl --context kind-microbank-advanced -n argo-rollouts rollout status \
  deploy/argo-rollouts --timeout=180s

Read the selected release's supported Kubernetes versions and CRD installation guidance. Investigate any apply error before deciding how to resolve a conflict.

2. Build a check with two fixed local targets

Save scripts/study/ledger_check.py. It permits only the two hardcoded in-cluster Service endpoints, disables proxies/redirects, validates case IDs, and performs bounded HTTP reads. It does not create transactions.

python
import argparse
import json
from pathlib import Path
import sys
import uuid
from urllib.request import HTTPRedirectHandler, ProxyHandler, Request, build_opener

TARGETS = {
    'stable': 'http://ledger.microbank.svc.cluster.local:8001',
    'candidate': 'http://ledger-canary.microbank.svc.cluster.local:8001',
}


class NoRedirect(HTTPRedirectHandler):
    def redirect_request(self, req, fp, code, msg, headers, newurl):
        raise ValueError('Unexpected redirect from local Ledger')


HTTP = build_opener(ProxyHandler({}), NoRedirect())


def read_json(url):
    if not any(url.startswith(base + '/v1/') for base in TARGETS.values()):
        raise ValueError('Only the declared local Ledger services are allowed')
    with HTTP.open(Request(url, headers={'Accept': 'application/json'}), timeout=3) as response:
        return json.load(response)


def verify(case, target, expected=None, fetch=read_json):
    base = TARGETS[target]
    account = str(uuid.UUID(case['account_id']))
    transaction = str(uuid.UUID(case['tx_id']))
    cents = case['expected_cents'] if expected is None else expected
    if type(cents) is not int or cents <= 0:
        raise ValueError('Expected cents must be a positive integer')
    balance = fetch(base + f'/v1/balances/{account}')
    entries = fetch(base + f'/v1/entries/{account}')
    if not isinstance(balance, dict) or not isinstance(entries, list):
        raise ValueError('Unexpected Ledger response shape')
    if not all(isinstance(entry, dict) for entry in entries):
        raise ValueError('Unexpected Ledger entry shape')
    matching = [entry for entry in entries if entry.get('txId') == transaction]
    if balance.get('balanceCents') != cents or len(matching) != 1:
        raise RuntimeError('Saved Ledger case did not match')
    return {'target': target, 'verified': True, 'matching_entries': 1}


def main():
    parser = argparse.ArgumentParser()
    parser.add_argument('--target', choices=TARGETS, required=True)
    parser.add_argument('--case', type=Path, default=Path('/case/case.json'))
    parser.add_argument('--expected-cents', type=int)
    args = parser.parse_args()
    try:
        print(json.dumps(verify(json.loads(args.case.read_text()), args.target, args.expected_cents)))
        return 0
    except (OSError, ValueError, KeyError, TypeError, RuntimeError) as error:
        print(f'Local Ledger check failed: {type(error).__name__}: {error}', file=sys.stderr)
        return 1


if __name__ == '__main__':
    raise SystemExit(main())

Save infra/study/Dockerfile.check:

dockerfile
FROM python:3.13-slim
WORKDIR /check
COPY scripts/study/ledger_check.py ./ledger_check.py
USER 65532:65532
ENTRYPOINT ["python", "/check/ledger_check.py"]

Record the resolved base-image digest. Build and load the small runner using a tag tied to the committed runner source:

bash
export CHECK_TAG="$(git rev-parse HEAD)"
docker build -f infra/study/Dockerfile.check -t "microbank-study/check:$CHECK_TAG" .
kind load docker-image --name microbank-advanced "microbank-study/check:$CHECK_TAG"
kubectl --context kind-microbank-advanced -n microbank create configmap ledger-check-case \
  --from-file=case.json=evidence/kind-case.json

Check that the working tree was clean when recording CHECK_TAG. The case contains synthetic IDs, not credentials; keep it private regardless. Before deployment, unit-test verify with fake responses for a matching case, wrong balance, duplicate entry, malformed response, and timeout. These tests check the verification function; you’ll test the running service separately below.

3. Generate the Rollout, analysis, and suspended schedule

Save scripts/study/canary_manifests.py. It reuses the actual baseline Ledger Pod template rather than guessing its environment variables. JSON generated here is ordinary Kubernetes input.

python
import copy
import json
import os
from pathlib import Path

root = Path('infra/study/k8s')
baseline = json.loads(Path('.local/canary-baseline.json').read_text())
items = baseline['items']
(root / 'recovery-apps.json').write_text(json.dumps(baseline, indent=2) + '\n')
ledger = next(item for item in items if item['kind'] == 'Deployment' and item['metadata']['name'] == 'ledger')
service = next(item for item in items if item['kind'] == 'Service' and item['metadata']['name'] == 'ledger')
image = 'microbank-study/check:' + os.environ['CHECK_TAG']


def check_job(target, expected=None):
    args = ['--target', target]
    if expected is not None:
        args += ['--expected-cents', expected]
    return {
        'backoffLimit': 0, 'activeDeadlineSeconds': 60, 'ttlSecondsAfterFinished': 3600,
        'template': {'spec': {
            'restartPolicy': 'Never', 'automountServiceAccountToken': False,
            'securityContext': {'runAsNonRoot': True, 'runAsUser': 65532, 'runAsGroup': 65532,
                                'seccompProfile': {'type': 'RuntimeDefault'}},
            'containers': [{'name': 'check', 'image': image, 'imagePullPolicy': 'IfNotPresent',
                            'args': args,
                            'securityContext': {'allowPrivilegeEscalation': False,
                                                'readOnlyRootFilesystem': True,
                                                'capabilities': {'drop': ['ALL']}},
                            'resources': {'requests': {'cpu': '25m', 'memory': '32Mi'},
                                          'limits': {'memory': '128Mi'}},
                            'volumeMounts': [{'name': 'case', 'mountPath': '/case', 'readOnly': True}]}],
            'volumes': [{'name': 'case', 'configMap': {'name': 'ledger-check-case'}}]
        }}
    }


analysis = {
    'apiVersion': 'argoproj.io/v1alpha1', 'kind': 'AnalysisTemplate',
    'metadata': {'name': 'ledger-candidate-check'},
    'spec': {'args': [{'name': 'expected-cents', 'value': '1000'}],
             'metrics': [{'name': 'saved-case', 'count': 1, 'failureLimit': 0,
                          'provider': {'job': {'spec': check_job('candidate', '{{args.expected-cents}}')}}}]}
}
rollout = {
    'apiVersion': 'argoproj.io/v1alpha1', 'kind': 'Rollout', 'metadata': {'name': 'ledger'},
    'spec': {'replicas': 2, 'revisionHistoryLimit': 2,
             'selector': copy.deepcopy(ledger['spec']['selector']),
             'template': copy.deepcopy(ledger['spec']['template']),
             'strategy': {'canary': {
                 'stableService': 'ledger', 'canaryService': 'ledger-canary',
                 'maxSurge': 1, 'maxUnavailable': 0,
                 'steps': [{'setWeight': 50}, {'pause': {}},
                           {'analysis': {'templates': [{'templateName': 'ledger-candidate-check'}]}},
                           {'setWeight': 100}]
             }}}
}
canary_service = copy.deepcopy(service)
canary_service['metadata']['name'] = 'ledger-canary'
cron = {
    'apiVersion': 'batch/v1', 'kind': 'CronJob', 'metadata': {'name': 'ledger-local-read'},
    'spec': {'schedule': '*/2 * * * *', 'suspend': True, 'concurrencyPolicy': 'Forbid',
             'startingDeadlineSeconds': 30, 'successfulJobsHistoryLimit': 2,
             'failedJobsHistoryLimit': 2, 'jobTemplate': {'spec': check_job('stable')}}
}
for filename, documents in [('canary.json', [service, canary_service, analysis, rollout]),
                            ('synthetic.json', [cron])]:
    (root / filename).write_text(json.dumps({'apiVersion': 'v1', 'kind': 'List',
                                            'items': documents}, indent=2) + '\n')

4. Establish stable before testing a candidate

Stop the Ledger port-forward before replacing its Deployment. Brief local downtime during this ownership handoff is expected. The original database/PVC remains.

bash
python3 scripts/study/canary_manifests.py
kubectl --context kind-microbank-advanced -n microbank delete deployment ledger
kubectl --context kind-microbank-advanced -n microbank apply -f infra/study/k8s/canary.json
kubectl argo rollouts --context kind-microbank-advanced -n microbank get rollout ledger

Wait for a healthy initial Rollout. The first creation establishes the stable revision; it does not run the old-versus-new canary gate. Inspect the ledger Service selector and the selected Pod hash. Apply the suspended CronJob, then trigger a single manual check:

bash
kubectl --context kind-microbank-advanced -n microbank apply -f infra/study/k8s/synthetic.json
kubectl --context kind-microbank-advanced -n microbank create job \
  --from=cronjob/ledger-local-read ledger-local-once
kubectl --context kind-microbank-advanced -n microbank wait --for=condition=complete \
  job/ledger-local-once --timeout=90s
kubectl --context kind-microbank-advanced -n microbank logs job/ledger-local-once

If it fails, inspect Job/Pod status and logs before proceeding. For another manual run use a new Job name. Once verified, unsuspend for a short observed window, then suspend again:

bash
kubectl --context kind-microbank-advanced -n microbank patch cronjob ledger-local-read \
  --type merge --patch '{"spec":{"suspend":false}}'
# Observe two scheduled runs locally, then stop scheduling.
kubectl --context kind-microbank-advanced -n microbank patch cronjob ledger-local-read \
  --type merge --patch '{"spec":{"suspend":true}}'

Do not run the second command immediately if you intend to observe scheduling. Check for active Jobs after suspension. Stable synthetic checks and candidate analysis share test logic but address different Services.

5. Prove the gate can fail

Edit generated canary.json: set the AnalysisTemplate's expected-cents default to 1001, and add a Pod-template annotation learning.learnwithsk.dev/revision: gate-negative-1 to the Rollout. Apply it. The annotation creates a new Pod-template revision using the same application image; this is a gate test, not an invented application regression.

Wait at the explicit pause. Inspect both Services' rollouts-pod-template-hash selectors, the Rollout's stable/current hashes, and ready endpoints. The candidate Service must select the new revision. With no traffic router, this lab does not assert a precise percentage of user requests; stable checks intentionally keep using the stable-only Service.

Resume the step, not all steps at once:

bash
kubectl argo rollouts --context kind-microbank-advanced -n microbank promote ledger
kubectl --context kind-microbank-advanced -n microbank get analysisruns,jobs
kubectl argo rollouts --context kind-microbank-advanced -n microbank get rollout ledger

The candidate check should exit nonzero because the actual saved balance is 1000. Inspect the specific AnalysisRun and its Job logs; expect failed analysis to block or abort progression. A Job that cannot start is missing evidence, not a successful negative assertion. Do not use promote --full to bypass the gate. Capture candidate selection before resuming: after an abort without traffic routing, the controller can point the canary Service back to stable. A post-abort selector alone does not identify what the failed Job tested.

6. Restore and prove the positive path

Restore the baseline Rollout Pod template from infra/study/k8s/recovery-apps.json, restore expected cents to 1000, and apply. Wait for the stable baseline to recover. Verify the saved case through stable. Then add a fresh annotation revision such as gate-positive-1, apply, inspect candidate selection at the pause, and resume one step. This time the AnalysisRun should succeed and the revision should promote.

A real application-image update uses the same sequence after building/loading the candidate image, but needs additional business and compatibility tests. The current read-only gate can miss a broken write path.

Evidence and return to Deployment mode

Save selected versions, baseline/candidate image IDs and hashes, both AnalysisRun outcomes, Job logs, schedule observations, and recovery verification. Keep failure evidence; TTL cleanup eventually removes Jobs.

Suspend the CronJob and wait for active checks to finish. Remove its CronJob, the Ledger Rollout, AnalysisTemplate, and candidate Service by their exact names. Restore the original stable Service selector to only app=ledger—remove the controller-added rollouts-pod-template-hash—then apply infra/study/k8s/recovery-apps.json to restore the Deployment. Inspect endpoints and restart the Ledger port-forward. Compare the restored declaration with the latest GitOps source before resuming that owner, so it cannot undo a hardening change. Only after the saved case and a fresh baseline check pass should you resume the former GitOps Application or start the failure-exercise module. Controller installation can remain idle until final cluster cleanup.

Explain the gate before claiming a release result

At the pause, identify stable and candidate by template hash, then identify the Service the Job will call. The wrong-1001 run should reach the saved case and reject its balance. An image-pull or DNS failure may also prevent promotion, but would not demonstrate that particular assertion. Keep the Job message with the AnalysisRun phase.

The subsequent correct-1000 run exercises the positive gate path. Because these revisions can use the same application image with different annotations, the result establishes the release machinery and read-check wiring. Fresh writes, schema compatibility, and concurrent consumer behavior remain outside this experiment. Complete the return-to-Deployment sequence before testing the next fault.

Sources

Argo Rollouts installation↗, Rollout specification↗, Job analysis↗, canary progression↗, and CronJob behavior↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.