Learning bite
Distributed data, retries, and idempotency
Trace ambiguous outcomes and distinguish a local database transaction from a distributed guarantee.
On this page
A missing response does not tell you the write failed
A caller submits a deposit; Accounts commits it; the connection closes before the caller receives a response. Retrying may be appropriate, but sending a new operation identifier can produce another deposit. An idempotency contract needs a defined key scope, payload consistency check, durable result, retention period, and behavior for concurrent requests.
A dictionary used in the DSA exercises is a useful local model, not shared durable storage. In a service, checking for a row and then inserting it leaves a race unless the storage constraints and transaction handling close it. A unique constraint can prevent duplicate rows, but the caller still needs a deliberate response when competing inserts collide.
Inspect the actual transaction boundaries
At the reference MicroBank revision, Accounts queries by account and idempotency key before inserting. The inspected transaction model does not declare a composite uniqueness constraint on that pair. Record this as a concurrency concern to investigate against the actual database schema; do not assume a sequential retry test proves concurrent safety.
The transaction and outbox row share a database commit. Publishing the pending outbox event and recording published_at are separate from SNS's acceptance. A crash between those steps can cause repeat publication. A consumer must handle redelivery; an outbox is not an exactly-once delivery promise.
Ledger's entity declares a unique transaction identifier, while its service publishes a settlement event inside a database transaction. Inspect whether a broker success followed by database rollback can expose a result that was not committed. Mark any proposed Ledger outbox or deduplication changes as design work until implemented and tested.
Use a failure matrix
| Failure point | What may be uncertain | Design question |
|---|---|---|
| Caller times out after Accounts commit | Whether acceptance occurred | Can the same key return the original result? |
| Publisher sends but does not mark the outbox row | Whether the next pass will resend | Can the consumer repeat safely? |
| Consumer commits but acknowledgement fails | Whether the message will return | Is duplicate processing durable and atomic? |
| Invalid message repeatedly fails | Whether progress is blocked or delayed | What are the retry limit, quarantine, and replay procedure? |
| Ledger cannot reach its database | Whether a request is merely pending or lost | Which durable records and queue evidence distinguish them? |
Trace one ambiguous publication
On paper, label four steps: Accounts commits transaction plus outbox; publisher sends the event; SNS accepts it; publisher commits the publication marker. Place a crash after the third step. The database still has an unmarked outbox row, but the broker may already have the event. Retrying publication is understandable; it can also cause duplicate delivery. The consumer must recognize the same logical operation durably, not merely remember it in one process's set.
Now move the crash to before the first commit. There may be no accepted transaction or outbox row. The same generic “timeout” symptom can therefore hide different persisted states. To reason about recovery, identify the last durable step and inspect it rather than assume a failed response means a failed write.
Budget retries and consistency
Choose an overall deadline, per-attempt timeout, maximum attempts, backoff, and jitter. Classify retryable failures and respect operation semantics. Retries at several layers can multiply traffic when a dependency is already struggling. A timeout can be ambiguous, so automatic retries for writes require the idempotency contract first.
Database isolation determines which concurrent behaviors a transaction permits. It does not make a database commit atomic with an external broker. Read-after-write expectations must also identify which service owns the read; eventual convergence is meaningful only with a completion definition, a timeout, and reconciliation when it does not happen.
Checkpoint and revision
Choose one failure row and draw a sequence with the crash between two explicit steps. Identify which durable data survive and how a retry is handled. Specify tests for a repeated key with a changed payload and two simultaneous submissions. Record these as source findings and proposed tests. They remain open concerns until you implement and verify a fix in MicroBank.
Checkpoint guide: a valid retry test checks a repeated key with the same payload, a changed payload with the same key, and competing requests. A database uniqueness constraint can reject a race, but the application must translate that result into the promised response. Keep these as proposed MicroBank tests until implemented; next, estimate how retries and backlog affect capacity.
Sources
Accounts transaction model↗, outbox publisher↗, and Ledger service↗. PostgreSQL isolation↗, SQS at-least-once delivery↗, and AWS retry with backoff↗.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.