Skip to content
← Advanced DevOps

Learning bite

Hash maps and sets for event analysis

Separate occurrence counting, event deduplication, and durable idempotency.

Documentation reviewed2026-10-01 · 3 min read
On this page

Define what a duplicate means

Two timeout records usually represent two observations. Two deliveries carrying the same event identifier may represent one business event. Deduplicating by message text can erase a genuine recurring failure; counting every delivery can exaggerate business activity.

Use a map for counts and a set for exact membership. In the following invented event fixture, identifiers are globally unique within the file. Real sources may need a composite key such as (producer, event_id). These records are not captured MicroBank logs.

Count occurrences without hiding repeats

Save as events.py:

python
from collections import Counter


def summarize(rows):
    deliveries = Counter()
    unique_events = Counter()
    seen = set()
    duplicates = 0
    for row in rows:
        event_id, kind = row["id"], row["kind"]
        if not isinstance(event_id, str) or not event_id.strip():
            raise ValueError("id must be a nonempty string")
        if not isinstance(kind, str) or not kind.strip():
            raise ValueError("kind must be a nonempty string")
        deliveries[kind] += 1
        if event_id in seen:
            duplicates += 1
            continue
        seen.add(event_id)
        unique_events[kind] += 1
    return dict(deliveries), dict(unique_events), duplicates


fixture = [
    {"id": "evt-a", "kind": "timeout"},
    {"id": "evt-a", "kind": "timeout"},
    {"id": "evt-b", "kind": "timeout"},
    {"id": "evt-c", "kind": "connection"},
]
assert summarize(fixture) == (
    {"timeout": 3, "connection": 1},
    {"timeout": 2, "connection": 1},
    1,
)

This preserves delivery counts while exposing repeated identifiers separately. Missing fields raise KeyError; invalid strings raise ValueError. For this file, the first occurrence of an identifier is authoritative. For real ingestion, explicitly reject an identifier reused with a different payload instead of silently treating it as a valid retry.

Explain each count in the fixture

Run python3 events.py in .local/advanced-design/; the assertions should pass silently. Walk the four rows yourself: the first evt-a adds one delivery and one unique timeout; the second adds a delivery and a duplicate but no unique event; evt-b adds a new timeout; evt-c adds a new connection event. This is why three timeout deliveries represent two distinct timeout events.

For a conflicting fixture, change the second evt-a to kind connection. The current code still calls it a duplicate and keeps the first occurrence's kind. That is documented behavior, not payload-conflict protection. A stricter version needs a map from event ID to the accepted payload (or a carefully specified representation), then rejects a changed payload under the same ID. Keep that extension separate from the baseline tests.

Bound the data you retain

For n deliveries, u kinds, and d distinct identifiers, expected work is O(n) and retained state is O(u + d), assuming bounded-size keys. Exact unbounded deduplication needs unbounded history. A retention window changes the promise: an older duplicate may be counted again.

A process-local set disappears after restart and is not shared across workers. It cannot establish durable transaction idempotency for MicroBank. That needs a database-backed contract and concurrency handling in the later design bites. A Bloom filter is also not an exact set: false positives matter if dropping a legitimate event is unacceptable.

Checkpoint and revision

Add an empty file, two different identifiers with the same kind, a blank identifier, and a conflicting repeated identifier. Describe the handling of each before changing the function. Explain why a ranking of frequent errors should usually count observations while a payment-like business operation needs a different duplicate policy.

Checkpoint answers: empty input gives empty counts and zero duplicates; two IDs with the same kind count as two unique events; a blank ID raises ValueError; a missing ID raises KeyError; conflicting reuse is not rejected by the baseline. Next, use an explicitly defined graph to reason about affected services rather than treating frequent logs as proof of dependency or cause.

Sources

Python Counter↗, MIT hashing↗, and SQS duplicate-delivery considerations↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.