Skip to content
← Advanced DevOps

Learning bite

Analysis conditions and failure handling

Make candidate evidence a real gate with explicit failure behavior.

Documentation reviewed2026-10-01 · 4 min read
On this page

Connect the check to the controller

Running a synthetic script beside a rollout does not automatically gate promotion. An AnalysisTemplate describes the measurement; an AnalysisRun evaluates it. The Rollout must reference it at the intended step.

The local lab uses a Job provider: a finite script verifies the saved case through the candidate-only Ledger Service. The controller uses successful Job completion as evidence that the check passed. A failed test exits nonzero. The Job has no retries and has a deadline so a hanging request cannot be mistaken for a pass.

Read the gate as a chain of evidence

An AnalysisTemplate is reusable configuration. An AnalysisRun is one evaluation made from it. The Job provider creates a finite Kubernetes Job; the test process exits zero only when its assertions pass. A container merely starting is not success, and retrying until something passes can hide intermittent failure unless the retry policy is explicit.

Use this invented decision table before creating resources:

Candidate observationWhat it establishesNext action
Correct saved transaction and balance, exit zeroThis candidate passed this read checkContinue only the intended next rollout step
Wrong balance, nonzero exitThe assertion failedStop progression and inspect application/test inputs
Image cannot pullThe check did not executeRepair test infrastructure; retain missing evidence
Request exceeds deadlineNo timely passing resultInspect target reachability and delay before retrying

The same image with an intentionally wrong expected balance is a useful negative gate test. It isolates the controller/test connection: you know the input assertion was changed. It cannot be reported as discovery of a broken application release.

Practice: imagine the stable endpoint passes while the candidate Job fails. Which result authorizes promotion? Neither a stable pass nor a green cluster overview can replace the candidate result. Inspect the AnalysisRun's provider message and Job logs, then verify the candidate Service selected the intended template hash at evaluation time.

A stable revision is established, a candidate pauses, and candidate-only analysis either permits the next step or stops progression for a mismatch or missing evidence; recovery verifies the saved case.

Open diagram at full size.

The split is the key decision: a completed passing assertion, a completed mismatch, and a test that could not run lead to different conclusions. The lab's intentionally wrong expected balance follows the mismatch branch.

Treat uncertainty explicitly

A failed test, a test that could not start, and a missing metric are different conditions. Inspect the AnalysisRun phase and message: Failed, Error, and Inconclusive are not successful promotion evidence. Exact transitions can depend on the controller version and provider; keep the selected release in the record.

If using Prometheus later, first verify that the query returns the expected number of series and enough recent samples. An empty vector, NaN, or stale data must not satisfy a success condition by accident. Do not average stable and candidate results together and call that a candidate gate.

A deliberate negative case

The lab initially verifies a real saved Ledger case. Then it changes the test's expected balance to an intentionally wrong value during a paused candidate rollout. The analysis should fail while the control-plane and test infrastructure remain available. This tests the gate itself; it does not claim that the application has a regression.

Checkpoint: show the AnalysisRun, Job exit result, selected candidate Pod hash, and the resulting Rollout state. Describe how you would distinguish a broken candidate from a broken test environment before retrying.

Do not infer traffic control from step percentages

A canary percentage does not by itself establish precise user traffic weighting. With no traffic-routing integration, replica proportions and Service selection affect behavior differently from a router-managed split. The MicroBank lab deliberately probes the candidate Service directly, so it tests that candidate's saved-case read result rather than claiming a measured fraction of real users saw it.

Record the analysis provider, target endpoint, selected revision, and result freshness. A successful stable request must not mask a failing candidate. The later platform handoff must preserve the latest tested configuration when changing workload ownership.

Checkpoint answer: a failed assertion has evidence that the test reached the application and found a mismatch; a start error lacks that evidence. Both block a confident pass but lead to different investigations. Preserve the negative result, restore the real expectation, and use a fresh candidate evaluation for the positive path described in recovery.

Sources

Analysis and experiments↗, Job analysis provider↗, and Prometheus analysis↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.