Learning bite
Capacity, overload, and recovery design
Estimate with explicit assumptions and turn bottleneck hypotheses into local measurements.
On this page
Start with quantities and units
Start capacity planning by making your assumptions visible. Separate arrival rate, completed work per second, latency, active concurrency, queue depth, stored bytes, and retention. A fast acceptance response can coexist with a growing Ledger backlog.
This arithmetic uses an invented worksheet fixture, not a MicroBank benchmark: arrivals are 12 events/second, successful processing is 10 events/second, and the difference lasts 30 seconds. With constant rates, no retries, and no other losses, the backlog increases by (12 - 10) × 30 = 60 events. If arrivals then stop and processing remains at 10/second, that extra backlog takes 6 seconds to clear. Ongoing arrivals, variable service time, failures, and pre-existing backlog change that answer.
For retained payloads, begin with events/second × bytes/event × retention_seconds. Database indexes, row metadata, replication, compression, and backups require separate estimates. Replace each fixture value with a measured or justified input before using the calculation for a decision.
Locate the limit before scaling
| Symptom | Possible constraint | Evidence to collect |
|---|---|---|
| Increasing journey completion time | Queue growth or a slow dependency | Acceptance and Ledger timestamps; backlog age |
| More replicas but no more completed work | Shared storage, locks, connection pool, broker | Per-stage timing and saturation |
| Repeated bursts after an outage | Synchronized retries | Attempt counts, retry timing, accepted unique operations |
| Good average latency but bad tail latency | Uneven work or contention | Distribution of real samples, not average alone |
| Host pressure during a lab | Too many local services or simultaneous builds | Host/engine memory and competing work |
A bounded queue needs a full-queue policy. Backpressure asks upstream to slow down; load shedding rejects selected work; concurrency limits bound in-flight work. Choose which work can wait or be rejected, and expose that choice to the caller. Make overload visible to callers, and preserve any work the service has already accepted.
Design recovery as a testable procedure
Record a recovery time objective (RTO: desired restoration time) and recovery point objective (RPO: tolerated data-loss window) as proposed requirements. Neither is an achieved result until measured. Replication, a restart, and a backup solve different problems. Plan a restore-and-verify exercise, including the data used to confirm recovery.
On the 16 GB RAM / 512 GB host, these opening calculations need only a document and small fixtures. In the later labs, run one local profile at a time with explicit resource limits. A one-node kind cluster and a two-node kubeadm lab do not prove high availability. Periodic synthetic users, load experiments, and deliberate faults remain local; they must not silently run against EKS/GKE or hosted CI.
Checkpoint and revision
Write a capacity worksheet with units, assumptions, and missing measurements. Choose one bottleneck hypothesis, one overload behavior, one user-oriented indicator, and a restore verification step. Connect them to the later SRE capacity bite and controlled failure lab.
Check the worksheet before the design lab
Using the invented backlog fixture, keep arrivals running at 8/second after the first burst while processing stays at 10/second. The spare processing rate is now 2/second, so the additional 60 events take 30 seconds to clear, not 6. If arrivals remain 12/second, the backlog does not clear under the stated model. This is why acceptance throughput and completed throughput must both appear in a capacity note.
For recovery, a proposed RPO of five minutes and RTO of thirty minutes means tolerated data loss and desired restoration duration, respectively. Neither can be inferred from “we take backups.” Identify the restore input, access to it, validation data, and measurements needed. Bring that worksheet into the MicroBank design lab before selecting more infrastructure.
Sources
Google SRE: handling overload↗, implementing SLOs↗, and AWS disaster-recovery objectives↗. The numerical worksheet is an original labelled fixture, not a result reported by these sources.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.