Skip to content
← Advanced DevOps

Learning bite

Scheduling and troubleshooting

Use evidence to locate a failure before changing a workload.

Documentation reviewed2026-10-01 · 3 min read
On this page

Diagnose from the outside inward

A user's failed transaction is a symptom. Separate scheduling, startup, routing, application, and dependency failures before choosing a fix.

ObservationFirst evidencePossible next step
PendingPod events, node allocatable resources, PVC stateCorrect an unsatisfied scheduling/storage constraint
ImagePullBackOffEvents and exact image referenceLoad the intended kind image or fix registry access
CrashLoopBackOffCurrent and previous container logsCorrect startup/configuration failure
Service unreachableEndpoints, selectors, readiness, DNSRepair the actual access path
API responds, Ledger result absentTransaction ID, outbox, queue, consumer logsLocate the asynchronous processing boundary

These are starting hypotheses, not diagnoses from status alone. OOMKilled is evidence of a memory termination; an exit code by itself needs context.

Diagnose an intentionally impossible placement

Use only the classroom fixture from the earlier bite, with its readiness path restored. Save a copy of .local/classroom.yaml. Add the following under the Deployment's spec.template.spec, beside containers:

yaml
nodeSelector:
  learning.learnwithsk.dev/room: unavailable

First inspect kubectl --context kind-microbank-advanced get nodes --show-labels and confirm no node has that label/value. Apply the changed classroom file. Inspect its Pods and describe the new Pending Pod. The expected scheduler event says the node selector cannot be satisfied; wording varies by version. A selector restricts eligible nodes. Resource requests, taints, affinity, and volume placement can restrict them too. A spare-looking CPU graph cannot override these conditions.

Remove the nodeSelector, apply again, and wait for the classroom rollout. Repeat the in-cluster HTTP check. The previous replica can remain available during the failed rolling update, so record which Pod was Pending and which one served the request. Restarting the node or rebuilding the image would not have corrected this deliberate mismatch.

Read a restart as a sequence

For a second paper exercise, suppose events show the image pulled and the container started, then the last termination reason says Error, the exit code is 1, and the previous log says a required setting is absent. Scheduling succeeded; image retrieval succeeded; application initialization failed. Inspect the setting's source and the container command next. CrashLoopBackOff describes delayed restart attempts, not the missing setting itself.

kubectl logs POD -c CONTAINER --previous retrieves the previous instance's log when available. describe pod supplies events and termination details even when the process emitted no useful log. Capture evidence before repeated restarts replace it. An init container that has not completed is another reason the main process may never have started.

Build a small evidence bundle

Once the module lab has deployed MicroBank, collect this bundle from its namespace. For the classroom exercise, use -n classroom and deployment/study-web instead.

bash
kubectl --context kind-microbank-advanced -n microbank get pods -o wide
kubectl --context kind-microbank-advanced -n microbank get events --sort-by=.lastTimestamp
kubectl --context kind-microbank-advanced -n microbank logs deploy/accounts --tail=80
kubectl --context kind-microbank-advanced -n microbank logs deploy/ledger --tail=80

For a restarted container, request --previous against its specific Pod and container. Logs may include test identifiers or SQL; redact them before publishing. Record timestamps and the failing image ID so you can compare observations across a deployment.

Reason before repairing

A successful HTTP health check does not clear the queue or commit a Ledger entry. A successful Pod restart does not establish why the previous process failed. Record the observed symptom, supporting evidence, hypothesis, smallest justified change, and verification result. If the evidence contradicts the hypothesis, revise it before making another change.

The module lab finishes only when a new synthetic deposit is verified through Ledger and the remaining application limitations are recorded.

Reasoning check: if the selector is repaired but the image cannot be fetched, you have advanced to a different failure, not disproved the first diagnosis. Follow events in order. Next, use the healthy classroom workload to reason about scaling and disruption before applying the full MicroBank manifests.

Sources

Debug applications↗, debug Pods↗, and resource troubleshooting↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.