Learning bite
Scheduling and troubleshooting
Use evidence to locate a failure before changing a workload.
On this page
Diagnose from the outside inward
A user's failed transaction is a symptom. Separate scheduling, startup, routing, application, and dependency failures before choosing a fix.
| Observation | First evidence | Possible next step |
|---|---|---|
| Pending | Pod events, node allocatable resources, PVC state | Correct an unsatisfied scheduling/storage constraint |
| ImagePullBackOff | Events and exact image reference | Load the intended kind image or fix registry access |
| CrashLoopBackOff | Current and previous container logs | Correct startup/configuration failure |
| Service unreachable | Endpoints, selectors, readiness, DNS | Repair the actual access path |
| API responds, Ledger result absent | Transaction ID, outbox, queue, consumer logs | Locate the asynchronous processing boundary |
These are starting hypotheses, not diagnoses from status alone. OOMKilled is evidence of a memory termination; an exit code by itself needs context.
Diagnose an intentionally impossible placement
Use only the classroom fixture from the earlier bite, with its readiness path restored. Save a copy of .local/classroom.yaml. Add the following under the Deployment's spec.template.spec, beside containers:
nodeSelector:
learning.learnwithsk.dev/room: unavailable
First inspect kubectl --context kind-microbank-advanced get nodes --show-labels and confirm no node has that label/value. Apply the changed classroom file. Inspect its Pods and describe the new Pending Pod. The expected scheduler event says the node selector cannot be satisfied; wording varies by version. A selector restricts eligible nodes. Resource requests, taints, affinity, and volume placement can restrict them too. A spare-looking CPU graph cannot override these conditions.
Remove the nodeSelector, apply again, and wait for the classroom rollout. Repeat the in-cluster HTTP check. The previous replica can remain available during the failed rolling update, so record which Pod was Pending and which one served the request. Restarting the node or rebuilding the image would not have corrected this deliberate mismatch.
Read a restart as a sequence
For a second paper exercise, suppose events show the image pulled and the container started, then the last termination reason says Error, the exit code is 1, and the previous log says a required setting is absent. Scheduling succeeded; image retrieval succeeded; application initialization failed. Inspect the setting's source and the container command next. CrashLoopBackOff describes delayed restart attempts, not the missing setting itself.
kubectl logs POD -c CONTAINER --previous retrieves the previous instance's log when available. describe pod supplies events and termination details even when the process emitted no useful log. Capture evidence before repeated restarts replace it. An init container that has not completed is another reason the main process may never have started.
Build a small evidence bundle
Once the module lab has deployed MicroBank, collect this bundle from its namespace. For the classroom exercise, use -n classroom and deployment/study-web instead.
kubectl --context kind-microbank-advanced -n microbank get pods -o wide
kubectl --context kind-microbank-advanced -n microbank get events --sort-by=.lastTimestamp
kubectl --context kind-microbank-advanced -n microbank logs deploy/accounts --tail=80
kubectl --context kind-microbank-advanced -n microbank logs deploy/ledger --tail=80
For a restarted container, request --previous against its specific Pod and container. Logs may include test identifiers or SQL; redact them before publishing. Record timestamps and the failing image ID so you can compare observations across a deployment.
Reason before repairing
A successful HTTP health check does not clear the queue or commit a Ledger entry. A successful Pod restart does not establish why the previous process failed. Record the observed symptom, supporting evidence, hypothesis, smallest justified change, and verification result. If the evidence contradicts the hypothesis, revise it before making another change.
The module lab finishes only when a new synthetic deposit is verified through Ledger and the remaining application limitations are recorded.
Reasoning check: if the selector is repaired but the image cannot be fetched, you have advanced to a different failure, not disproved the first diagnosis. Follow events in order. Next, use the healthy classroom workload to reason about scaling and disruption before applying the full MicroBank manifests.
Sources
Debug applications↗, debug Pods↗, and resource troubleshooting↗.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.