Learning bite
Toil, capacity, and load
Measure a bottleneck and automate a repeated task with a clear boundary.
On this page
Distinguish repetitive work from useful investigation
Toil is recurring operational effort that grows with the service and can often be reduced. Repeatedly collecting the same diagnostics is a candidate for automation. Investigating a new failure is not automatically toil.
For MicroBank, start with a script that gathers bounded Pod status, events, image identities, and recent logs from the named local namespace. Keep it read-only and redact sensitive values. The script should fail clearly if the expected context is unavailable.
Decide whether automation buys useful time
Use an invented weekly worksheet: gathering the same local diagnostics takes 4 minutes and happens 10 times, or 40 minutes. Building a dependable collection script takes 3 hours and maintenance takes 5 minutes per week. Ignoring other costs, the net saving is 35 minutes per week, so the initial effort is recovered in about 5.1 weeks. That calculation does not decide the task alone: fewer omissions and safer context selection may matter, while a rarely used fragile script can cost more than it saves.
Start with read-only collection and a named context. Test the wrong-context case, missing tool, and partial command failure. Avoid adding automatic restarts before you can describe when they help and when they would erase evidence. Automating an uncertain diagnosis only repeats the uncertainty faster.
Capacity work asks a related question: how does completed work change as demand changes? A load test checks a chosen expected workload, a stress test seeks limits, and a soak test observes behavior over a longer period. This course begins with a small sequential sample so you can explain the measurement before adding concurrency or duration.
Run a small capacity experiment
Choose one question, such as whether sequential deposits are limited by Accounts response time or Ledger convergence. Record idle host/engine memory, then run a small bounded series of fresh synthetic journeys. Keep concurrency at one initially. Measure request failures, elapsed duration, queue behavior, and resource pressure.
Keep periodic canary checks separate from load tests. Load tests intentionally vary arrival rate/concurrency; synthetic checks usually sample a small workflow. A Mac under build load can distort both. Record competing workloads and stop if the host becomes unresponsive, memory pressure rises sharply, or the experiment threatens unrelated work.
Interpret the result narrowly
One local run cannot predict an EKS node count or cloud cost. A memory limit that prevents host exhaustion can also cause OOM failures. Increasing it may mask a leak; decreasing it may create a misleading bottleneck.
Checkpoint: publish a table of measured conditions and results, one bottleneck hypothesis, and a repeatable next experiment. Keep the host's 16 GB RAM and remaining disk space in the record. Report only the throughput you measured, and keep production SLA decisions separate from this short local run.
Worked interpretation: in an invented run, ten journeys take 50 seconds, with two failures and eight verified results. Attempt rate is 0.2/second and verified completion rate is 0.16/second across that interval. Reporting “ten successful transactions” would be wrong. The slowest stage and waiting time still need measurement before choosing a fix. Next, use the SRE lab to calculate labelled fixtures and record your own small sample separately.
Sources
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.