Skip to content
← Advanced DevOps

Learning bite

Incidents and postmortems

Write a useful recovery record without inventing an outage.

Documentation reviewed2026-10-01 · 3 min read
On this page

Separate restoring service from explaining the cause

An incident is a period when service behavior requires coordinated response. Detection identifies an observed problem; mitigation reduces its impact; recovery verification checks that the user outcome works again. A postmortem explains what contributed and what should change afterwards. You can mitigate before knowing every cause, but should not turn a successful mitigation into an unsupported causal claim.

For this invented paper timeline, a saved-case probe fails at 10:00; at 10:01 you find its Ledger port-forward stopped; at 10:02 you restart the forwarding process; at 10:03 the same case passes. This supports an observer-access interruption. It does not establish that Ledger stopped, that the queue lost messages, or that a database restarted. Detection was at 10:00, but the actual failure start may be unknown. Say so rather than inventing a duration.

In a team response, distinguish the person coordinating decisions, the operator investigating/mitigating, and the person recording/communicating updates. In a solo lab you perform those jobs sequentially. Keep one timeline and pause unrelated changes so your own experiments do not obscure the evidence.

Practise incident handling locally

During a scheduled MicroBank fault exercise, record when the symptom began, when the check detected it, your working hypothesis, what you changed, and when verification recovered. Label the record a lab incident. Keep the distinction between a lab incident and production experience clear.

Separate observations from hypotheses. “Ledger replicas were zero at 14:02” is an observation. “The queue lost data” needs evidence. Capture current events and bounded logs before repeated restarts erase useful context.

Choose the next observation

Use this decision tree for the local saved-case check. It is a diagnostic order, not a claim that only these failures exist:

text
Saved-case check failed
├─ Did the check start and complete?
│  └─ No: inspect runner, schedule, deadline, and telemetry freshness
└─ Yes: was its endpoint reachable?
   ├─ No: inspect port-forward / Service / endpoints / network path
   └─ Yes: did the response match the saved transaction and balance?
      ├─ No: inspect case, revision, persisted rows, and application logs
      └─ Yes: inspect the assertion and reporting path

Try the 10:00 example through the tree before reading its conclusion again. A failed observer path takes you to connectivity evidence first; it does not justify deleting data or injecting another fault.

Use a compact incident record

markdown
## Local incident record
- Scope and environment:
- Source/image revision:
- User journey affected:
- Start / detection / mitigation / verified recovery:
- Observations with timestamps:
- Hypotheses considered and rejected:
- Change that restored service:
- Data verification and remaining uncertainty:
- Follow-up action, owner, and verification method:

A recovery command succeeding is not the recovery criterion. Rerun the saved transaction case and inspect the matching Ledger entry. If the case cannot be reconstructed, say so instead of creating a fresh one and claiming the original transaction recovered.

Learn from the mechanism

A useful postmortem explains the causal chain and contributing conditions: for example, a queue consumer was unavailable, the process-health check stayed green elsewhere, and no end-to-end signal was observed. Improving the probe addresses a different gap from adding a replica.

Checkpoint: choose one follow-up that could prevent recurrence and one that could shorten diagnosis. Give each a test. Avoid blaming a person or listing “be more careful” as the only corrective action. A short record with clear observations and concrete follow-ups is enough.

Checkpoint answer: “restart the port-forward automatically and expose a separate observer-health signal” is testable; “be more careful” is not a mechanism. A prevention action and a diagnosis action can both be useful, but give each a specific owner and retest. Next, decide which recurring response work is worth automating.

Sources

Postmortem culture↗, incident management↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.