Learning path / 03
Advanced DevOps
Build DSA, system design, PostgreSQL, and SQL skills, then operate MicroBank with Kubernetes, GitOps, observability, and recovery exercises.
Before you begin
- Python functions, collections, exceptions, and basic tests from DevOps Foundations.
- Repeatable builds, container networking, Terraform, configuration management, and CI/CD concepts.
- Work through one local lab profile at a time and measure its resource use.
Your learning progress
0 of 59 available items completed
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.
The learning sequence
Explore a module to see its bites and proposed checkpoint.
01DSA and system design for DevOpsReason about algorithms and MicroBank failure boundaries before operating the cluster.Available
- DSA: problem contracts and complexityLearning bite
- Hash maps and sets for event analysisLearning bite
- Graphs, BFS, and DFS for dependency questionsLearning bite
- Topological ordering and cycle detectionLearning bite
- Heaps and top-k operational summariesLearning bite
- Binary search and timestamp boundariesLearning bite
- Sliding windows and rate-limit boundariesLearning bite
- Lab: test an operational analysis toolkitLab guide
- Practice: 15 algorithm problems and operational follow-upsCheckpoint
- System design: requirements, interfaces, and trade-offsLearning bite
- Distributed data, retries, and idempotencyLearning bite
- Capacity, overload, and recovery designLearning bite
- Lab: review MicroBank before KubernetesLab guide
Run the offline algorithm tests and produce a source-backed MicroBank design review; use the optional 15-problem practice set for further depth.
02Database foundations with PostgreSQLLearn SQL and how to operate MicroBank’s databases before moving to Kubernetes.Available
- PostgreSQL and your first databaseLearning bite
- Tables, data types, and constraintsLearning bite
- SQL reads, writes, joins, and aggregatesLearning bite
- Transactions, isolation, and lock waitsLearning bite
- Indexes and query plansLearning bite
- Database roles and connection budgetsLearning bite
- PostgreSQL diagnosis and maintenanceLearning bite
- Schema migrations, backups, and restore boundariesLearning bite
- Lab: practise SQL and recovery with PostgreSQLLab guide
- Lab: inspect MicroBank databases and plan their operationLab guide
Verify a disposable SQL fixture and logical restore, then inspect the real MicroBank schemas and document their operating risks.
03Kubernetes foundations and operationsUnderstand workloads and reconciliation before adding delivery controllers.Available
- Architecture and kindLearning bite
- Deployments, Services, and configurationLearning bite
- Storage, probes, and resource requestsLearning bite
- Two-node kubeadm operationsLearning bite
- Scheduling and troubleshootingLearning bite
- Scaling and disruption boundariesLearning bite
- Lab: run the MicroBank baseline in kindLab guide
Explain a workload failure using events, logs, probes, and resource observations.
04GitOps foundationsSeparate artifact creation from desired-state reconciliation.Available
- Kustomize bases and initial overlaysLearning bite
- Argo CD applications and syncLearning bite
- Drift, pruning, and recoveryLearning bite
- Delivery pipeline feedback and artifact promotionLearning bite
- Lab: reconcile MicroBank from a local overlayLab guide
Make a Git change, inspect the rendered configuration, and explain the resulting cluster change.
05Monitoring and observabilityUse telemetry to answer a specific operating question.Available
- Prometheus metrics and GrafanaLearning bite
- Alertmanager and actionable alertsLearning bite
- Logs with Loki and AlloyLearning bite
- Tracing with OpenTelemetryLearning bite
- Lab: observe the local MicroBank journeyLab guide
Diagnose a controlled failure by connecting an observed symptom to supporting telemetry.
06DevSecOps implementationIntroduce security checks throughout delivery and operations.Available
- Threat modelling and secretsLearning bite
- Image and IaC scanningLearning bite
- SBOMs, signatures, and provenanceLearning bite
- Identity and workload hardeningLearning bite
- Lab: remediate and verify one security findingLab guide
Record a finding, assess its impact, and verify a targeted remediation.
07Site reliability engineeringUse user-facing indicators to guide operational decisions.Available
- SLIs and justified SLOsLearning bite
- Error budgets and alertingLearning bite
- Incidents and postmortemsLearning bite
- Toil, capacity, and loadLearning bite
- Lab: define an SLI and write a recovery recordLab guide
Define a measurable service indicator and use a local incident to evaluate its usefulness.
08Canary releases and synthetic usersConnect release analysis to repeatable local user-journey checks.Available
- Argo Rollouts canary stepsLearning bite
- Local periodic synthetic journeysLearning bite
- Analysis conditions and failure handlingLearning bite
- Recovery and release evidenceLearning bite
- Lab: gate a local Ledger canary with synthetic checksLab guide
Demonstrate a candidate-specific gate that blocks promotion on a failed check, then verify recovery.
09Controlled failure exercisesTest an explicit hypothesis in disposable local preproduction.Available
- Baseline and blast radiusLearning bite
- One reversible fault at a timeLearning bite
- Observe, recover, and clean upLearning bite
- Lab: interrupt Ledger, recover, and explain the evidenceLab guide
Record the hypothesis, actual behaviour, recovery, and limits of the experiment.
Continue the working MicroBank project
Advanced DevOps has 47 learning bites, eleven practical labs, and one optional practice checkpoint across nine modules. Start after completing the DevOps Foundations MicroBank implementation, including a passing local transaction check. Each bite explains a concept, gives an exercise or inspection task, and ends with a checkpoint and primary references.
Begin with DSA and system design: seven algorithm bites, an offline Python lab, an optional 15-problem practice set, three system-design bites, and a MicroBank design review. The opening exercises need no running cluster or cloud resources. You can return to the optional checkpoint at any time; the design and deployment labs do not depend on it.
Next, Database foundations with PostgreSQL introduces SQL, constraints, transactions, indexes, roles, connection budgets, diagnosis, migrations, and recovery. Eight bites and two labs use a small isolated PostgreSQL 15 fixture, then inspect MicroBank's actual Accounts and Ledger databases. No prior SQL is required. Keep only one local runtime profile active.
You’ll then move the local container application to Kubernetes: deploy it, reconcile changes, measure a user journey, review security, check a release candidate, and recover from one controlled failure.
Build confidence one small result at a time
Each module has a different first step. Algorithms begin with deterministic files and hand traces. PostgreSQL begins with one isolated schema and queries whose expected rows are shown. Kubernetes begins with a small classroom web page: read its objects, replace a Pod, make a readiness path fail, diagnose an impossible selector, and restore it. Only then package MicroBank. Keep the separate kubeadm VM exercise separate from the kind application profile.
For each worked example, make a prediction before running or reading its answer. Explain the result in your own words, change one input, and predict again. The answer guides are there to reveal a mistaken model early. A passing command is useful when you can also say what it established and what it did not test.
GitOps adds an offline render before a controller, observability adds sample interpretation before a dashboard, and the release/failure modules add paper decisions before controlled local execution. Keep the examples and later lab evidence distinct: a fixture value is an expected answer, while your operating record contains the outcome you actually observed.
The practical progression
| Module | Bites | Practical evidence to collect |
|---|---|---|
| DSA and system design | 10 | Test an operational analysis toolkit and review the MicroBank design; optional algorithm practice |
| Database foundations with PostgreSQL | 8 | Verify SQL, lock behavior, and a logical restore, then inspect MicroBank data ownership |
| Kubernetes | 6 | Run MicroBank in kind, retain data, and verify a real deposit |
| GitOps | 4 | Render a Kustomize overlay and reconcile it with Argo CD |
| Observability | 4 | Export a local journey check, graph its freshness, and observe an alert |
| DevSecOps | 4 | Investigate and remediate an actual finding, then verify the artifact and application |
| SRE | 4 | Define an SLI, calculate a labelled fixture, and record real observations |
| Progressive delivery | 4 | Schedule local synthetic reads and gate a Ledger candidate with Argo Rollouts |
| Failure exercises | 3 | Interrupt Ledger, restore it, and verify the original transaction |
Keep the local lab manageable
Start with one disposable kind node. Use a separate kubeadm exercise with one control-plane VM and one worker to learn node administration; that topology is not highly available. Run these tracks separately. On the 16 GB RAM / 512 GB storage host, build images sequentially, inspect remaining storage, and measure host and container-engine pressure before adding controllers or telemetry.
Prometheus and Grafana are the first observability exercise. Alloy/Loki and OpenTelemetry instrumentation are separate additions guided by their bites; you do not need to run the full stack together. Stop optional tools when moving to a memory-heavier canary exercise. Version tags and image digests must be selected from current official release documentation and recorded with each run.
Periodic synthetic checks and deliberate faults run only locally. The CronJob begins suspended and belongs only to the local exercise. A read-only saved-case check, a fresh deposit journey, and a candidate-specific rollout gate establish different things. Cloud profiles and hosted CI must not inherit this scheduler.
Be precise about what MicroBank demonstrates
The MicroBank revision↗ defines the application behavior used in these lessons. Accounts' process-health response does not prove Ledger convergence. The lessons call out its backend authorization and settlement-consumer gaps. Use actual transaction records to verify results; the frontend includes mock values. Observability endpoints and rollout analysis are additions in these lessons, not features presumed to exist upstream.
The canary lab checks Ledger's read path against a real saved case. Stable and candidate share local dependencies, so it does not certify concurrent consumer behavior or schema compatibility. Its deliberate wrong-expectation run tests whether the gate blocks promotion; it is labelled as a gate test rather than an invented application regression.
Lessons are marked Reviewed, not Lab verified. Source review, snippet checks, and a successful site build do not establish a completed deployment. Record your actual environment, selected versions, results, and failures before treating a lab as executed.
Additional roadmap coverage
Applicable topics from Week 1↗ and Week 4↗ of the AI Platform Engineering Handbook strengthen this path with original scaling/disruption and delivery-feedback bites. Existing lessons now connect golden signals to actual MicroBank evidence, exercise supply-chain verification failures, and clarify canary traffic and identity boundaries.
Before the canary handoff, review and save a snapshot so recovery preserves your hardening changes. The later platform labs establish a canonical base and a single GitOps owner. These are guided procedures; a successful site build still does not prove a complete MicroBank runtime.
Source and adaptation
The opening module uses the two supplied DSA interview articles as topic inspiration, with newly authored explanations and exercises checked against Python documentation, MIT algorithm material, and primary system-design references. LeetCode links point to the original problem pages; no problem statements or solution listings are copied. Company interview-frequency claims and broad infrastructure analogies from the articles are not treated as established facts.
The database module is newly authored against version-matched PostgreSQL primary documentation and inspected MicroBank source. Its teaching schema is separate from the application; it does not claim to implement a production ledger, database HA, PITR, or application migrations. Each database bite links the relevant PostgreSQL reference.
The topic structure draws on the collaborative Compute Central curriculum↗, adapted with permission. The explanations and MicroBank exercises are written for this path, and each lesson links its primary technical references.
Core coverage and optional depth
The core path follows Kubernetes from architecture and objects through workloads, Services, configuration, storage, probes, scheduling, and basic operations. It then applies Kustomize/Argo CD, metrics and alerting, security review, SRE decisions, a local read-path canary, and one reversible failure. Logs and tracing start with interpreted examples and offer separate integration work so they do not require a simultaneous large stack.
The wider source curriculum also covers more than this path implements: Helm chart authoring, detailed CNI/CSI internals, Gateway/Ingress controllers, advanced StatefulSet and DaemonSet operations, multi-tenant quotas, etcd recovery and cluster upgrades, node autoscaling, full Vault operations, service meshes, compliance programs, and organizational on-call practice. Some appear as interview or design extensions; they are not completed labs here. Choose them after the relevant foundations, with their own prerequisites and verification. The custom DSA and PostgreSQL modules use their linked primary references rather than claiming equivalent chapters in the source curriculum.
What comes next
Platform engineering continues with 27 bites and six labs on advanced GitOps, environment promotion, Kyverno, EKS/GKE adapters, cost controls, and a Backstage or Port IDP. OpenShift, service meshes, deeper Ansible, and a three-VM Nomad comparison remain optional depth. Follow implementation evidence in Build, Break & Operate.
Interview prep arena
Use these questions to practise explaining decisions with evidence. Explain the request or control path, name the observation that would test your hypothesis, propose the smallest justified action, and verify the user outcome. Answer from your actual MicroBank revision or a labelled fixture. Do not invent a production incident or promise zero downtime from a checklist.
Choose a round: Kubernetes · Disruption and delivery · Security · Application diagnosis · Reliability · Algorithms · Project evidence.
Kubernetes control and request paths
Revisit architecture, workloads and Services, scheduling, and probes and storage.
ADV-01. Browser-to-pod request flow
Trace a browser request to a Kubernetes pod. Where do load balancers, reverse proxies, and Ingress fit? Mark DNS, TLS termination, routing, Service selection, the actual service-network implementation, and the application listener. Draw only components your deployment uses; Ingress is not mandatory for every request. Follow the return path and identify an observation at each hop.
ADV-02. What kubectl apply does
What happens between kubectl apply and a usable application? Distinguish client-side and server-side apply, API authentication/authorization, admission, persistence, and asynchronous controller reconciliation. Explain field ownership where relevant. An accepted API change does not mean the scheduler, kubelet, readiness checks, and traffic path have all succeeded. See server-side apply↗.
ADV-03. Scheduling and Pending pods
How is a pod assigned to a node, and how do you debug one that remains Pending? Follow scheduler filtering, scoring, and binding. Inspect events, resource requests, taints/tolerations, affinity, topology, and storage constraints; distinguish an unscheduled pod from one waiting for image or volume setup. Use the scheduler documentation↗ to explain why scheduling considers more than current CPU usage.
ADV-04. Workload controllers and self-healing
When would you use ReplicaSet, Deployment, StatefulSet, DaemonSet, Job, or CronJob? How do they respond to failure? Explain desired replicas, rollout history, stable identity/storage, per-node work, and finite tasks. Distinguish a kubelet restarting a container from a controller replacing a pod. Replacement alone cannot repair corrupt data or a consistently failing configuration.
ADV-05. Readiness, liveness, and startup
How do these probes determine whether a pod receives traffic or a container restarts? Explain why a Running pod may not be Ready, how startup probing changes the timing of other probes, and why an overbroad liveness dependency can cause restart cascades. Define a meaningful MicroBank check; process health does not establish transaction completion. See probe concepts↗.
ADV-06. Service discovery and pod communication
How do pods discover and communicate with a Service and with each other? Distinguish DNS naming, ClusterIP, headless Services, EndpointSlices, CNI networking, and the installed proxy or eBPF dataplane. Trace cross-namespace names and direct pod traffic. DNS resolution is only one step; policy, routing, ports, and readiness still matter.
ADV-07. CrashLoopBackOff without useful logs
How do you diagnose repeated restarts when current logs show no errors? Inspect the affected container, previous logs, termination reason and exit status, events, startup command, probes, configuration, and resource limits. Account for init containers, OOM kills, and logs written somewhere other than stdout. Use an appropriate diagnostic container if needed, while preserving evidence. Pod debugging↗ explains available inspection methods.
ADV-08. Running pods and 502, 503, or 504 responses
Where is the error generated, and how would your investigation differ for each status? Start with a failing request and its time/request ID. Compare load-balancer, Ingress/proxy, and application evidence; inspect Service selectors, EndpointSlice addresses/readiness, port/targetPort, protocol/TLS, and the upstream listener. A 502 suggests an invalid upstream response, a 503 unavailability, and a 504 an upstream timeout, but the emitting component determines the diagnosis. Check dependencies and timeout budgets. Do not begin with an unexplained restart or scale-up. Revisit Service debugging↗ and HTTP status semantics↗.
ADV-09. Successful deployment, old version
Deployment reports success, but users still reach the old application. Where do you start? Compare the expected revision/digest with the actual pod image IDs and rollout status, then trace traffic selectors, Ingress backend, rollout weights, target groups, and the environment URL. Check mutable tags, missed template changes, stale caches, or a skipped deployment. If the pipeline selected the wrong target, use DEV-09.
ADV-10. A node becomes NotReady
How do you distinguish lost node heartbeats, kubelet failure, runtime failure, network failure, and resource pressure? Correlate conditions, events, node reachability, kubelet/runtime logs, certificates, CNI, and disk availability. Explain where workloads can safely move before proposing cordon/drain, and what evidence would justify restoring or replacing the node.
ADV-11. Kubelet keeps restarting
How would you isolate persistent kubelet restarts? Review service exit information, configuration changes, runtime connectivity, certificates, resource pressure, and version compatibility. Compare an affected node with a healthy one. Stabilize or isolate one node at a time; restarting every node would erase the useful comparison and widen the incident.
ADV-12. StatefulSet pod will not return
A deleted StatefulSet pod fails to recreate correctly. How do you investigate without losing data? Separate a terminating old pod, controller ordering, readiness, PVC binding, CSI attachment, node availability, access modes, and volume topology. Inspect ownership and storage state before changing anything. Do not delete PVCs or force-delete a potentially still-running writer as a routine fix. Consult StatefulSet behavior↗ and the installed storage driver's recovery procedure.
Disruption, scaling, and release safety
Revisit scaling, kubeadm, canary delivery, and local jobs.
ADV-13. HPA and horizontal versus vertical scaling
How does HPA decide to change replica count? When would horizontal or vertical scaling help? Identify the metric source, target, resource requests for utilization-based scaling, stabilization, missing metrics, and capacity constraints. Explain what scaling cannot fix, such as lock contention or a shared database limit. Read the HPA algorithm↗.
ADV-14. Pending pods do not trigger node scale-up
What would you check when cluster autoscaling adds no nodes? First establish why the pods cannot schedule. Review autoscaler events and access, node-group limits, quotas, compatible instance capacity, affinity/taints, and whether a new node could satisfy storage and resource constraints. Differentiate node provisioning from HPA replica decisions. This is a cloud/cluster design extension; the local kind lab does not supply managed node autoscaling. See node autoscaling↗.
ADV-15. Cross-namespace NetworkPolicy
How would you permit only the required service-to-service path and debug an accidental block? Identify source/destination pod and namespace labels, ingress and egress rules, ports, and DNS needs. Verify that the CNI enforces NetworkPolicy and test both an allowed and a denied flow. Policies are additive; reason about all policies selecting each endpoint. Use the NetworkPolicy specification↗.
ADV-16. TCP and UDP on the same port
How would you expose both protocols for one service? Describe separately named Service ports with explicit protocols, then verify the selected load-balancer/provider supports the combination. Standard Ingress handles HTTP(S), not arbitrary TCP/UDP; other controller-specific or supported gateway/load-balancer mechanisms need explicit review. See Service protocol support↗ and Ingress boundaries↗.
ADV-17. Node-pressure eviction and QoS
A critical pod is evicted. What led to it, and how would you reduce recurrence? Check the pressured resource, requests versus usage, priority, QoS, node reservations, ephemeral storage, and capacity. QoS influences behavior but is not immunity from eviction. PDBs do not prevent node-pressure eviction. Verify your explanation against eviction behavior↗.
ADV-18. Container logs overload node disks
How do you diagnose and limit log-driven disk I/O? Identify high-volume containers and collectors, rotation and retention settings, filesystem capacity, and backpressure. Balance sampling/verbosity and collection buffering with the evidence you need. State what is configured by the kubelet/runtime versus the logging backend. See Kubernetes logging architecture↗.
ADV-19. Rolling updates and stateful availability
What happens during a rolling update, and why can it still cause downtime? How would you design a stateful upgrade? Account for surge/unavailable limits, spare capacity, readiness, connection draining, preStop, grace periods, sessions, and backward-compatible expand/contract database migrations. Discuss replication/quorum and rollback compatibility. PDBs constrain supported voluntary evictions; workload rolling updates use their own rules. Define and measure the availability objective rather than guarantee zero downtime. See Deployments↗ and disruptions↗.
ADV-20. Blue-green or canary
When would you choose blue-green over canary, and how would you verify promotion or rollback? Compare traffic control, capacity overhead, data compatibility, observation time, and exposure. Explain what candidate-specific checks actually cover. The local MicroBank read-path gate does not establish safe schema migration or fresh-write correctness for every release.
ADV-21. Cluster upgrade without user-visible interruption
How would you plan and execute a cluster upgrade? Inventory API removals, version skew, add-on/CSI/CNI compatibility, backups and restore procedure, workload redundancy, and spare capacity. Sequence control-plane and node changes using the distribution's supported procedure, cordon/drain within disruption constraints, and verify real traffic after each step. The two-node kubeadm exercise is not an HA cluster. Consult the version-specific upgrade procedure↗.
ADV-22. Production DNS breaks after an upgrade
Staging passed, but production CoreDNS fails. How would you restore service and plan a patch? Pause further rollout, compare versions and configuration, inspect DNS pods, Service/EndpointSlices, upstream resolvers, CNI/network policy, and kubelet DNS settings. Choose a supported component recovery using known-good evidence; do not assume an in-place cluster downgrade is safe. Verify both cluster-service and external lookup paths. See DNS debugging↗.
Workload security
Revisit threats and secrets, scanning, provenance, and workload identity.
ADV-23. Configuration and secrets across environments
How do ConfigMaps and Secrets differ, and how would you supply and rotate them across environments? Explain access, storage protection, application reload behavior, and references in rendered manifests. Encoding is not encryption. Keep plaintext credentials out of Git, images, logs, and unrestricted state. For fifty-service secret delivery, use PLT-03; Terraform persistence has a separate PLT-20 question.
ADV-24. Secure an image before release
Which controls belong before an image is promoted? Discuss base selection and updates, dependency and image scans, SAST and appropriately scoped DAST, secret checks, SBOMs, signature/provenance verification, and exception ownership. Distinguish what each check observes; a passed scan is time-bound evidence, not proof that an application cannot be exploited.
ADV-25. Scanned image, exploited workload
What would you investigate after a scanned image is exploited? Preserve evidence and assess credentials and reachability, application flaws, runtime configuration, privileged access, RBAC, seccomp/capabilities, writable filesystems, and monitoring gaps. Choose containment based on the actual exposure. Explain which runtime protections and credential changes address the finding without asserting that one hardening flag would have prevented every exploit.
ADV-26. Kubernetes RBAC
How would you implement and verify least-privilege access? Separate user and workload identities, namespaced and cluster-scoped permissions, RoleBindings, and escalation-sensitive actions. Test an allowed operation and a denied operation. Explain why permission to create workloads or read Secrets can carry broader consequences than its name suggests.
Application and data diagnosis
Revisit SQL queries, transactions and locks, connection budgets, database diagnosis, distributed data, capacity, and tracing.
ADV-27. Latency jumps with normal CPU and memory
An API becomes 8–10 times slower. Where do you start? Scope affected routes/users and latency percentiles, then split request time across queues, connection pools, DNS/TLS, locks, database queries, storage, and external dependencies. Compare errors, retries, saturation, changes, and a known-good interval. Healthy host averages do not rule out a dependency or per-request bottleneck.
ADV-28. Regression only at peak traffic
A new release fails only at peak traffic and rollback helps. What do you investigate next? Compare old/new query behavior, pools, concurrency, contention, memory allocation, retries, and resource limits under a bounded representative workload. Preserve the artifact and timing evidence; rollback is a useful observation, not proof of a single cause. Reproduce locally within an explicit stop condition.
ADV-29. Database connection errors at peak
How would you separate pool exhaustion, leaked connections, database limits, authentication, and network faults? Inspect acquisition wait time, active/idle connections, transaction duration, query locks, retry volume, and limits per instance and across replicas. Explain why adding application replicas may worsen the database connection budget.
ADV-30. HTTP 200 but the change is missing
Where do you look when an API returns success but a user cannot see the update? Establish the endpoint's success contract and trace the write, transaction commit, outbox/queue, consumer, read model, replica, cache, and frontend. Use a correlation or transaction ID. For MicroBank, distinguish an accepted deposit from observed Ledger convergence and durable correctness.
ADV-31. Cache serves outdated information
How would you diagnose stale results after introducing caching? Inspect cache keys, TTLs, invalidation, tenant/user scope, write ordering, replica lag, browser/CDN behavior, and concurrent refreshes. Separate intentional bounded staleness from an incorrect key or missing invalidation. Choose a consistency requirement before proposing shorter TTLs everywhere.
ADV-32. One slow dependency stalls the application
Which resilience mechanisms would you consider, and what could go wrong with each? Compare timeout budgets, bounded concurrency/bulkheads, circuit breakers, retry limits with jitter, queues, and honest degraded responses. Require idempotency where retries can duplicate effects. Unbounded retries or fabricated successful responses can amplify the failure.
ADV-33. Scheduled work executes more than once
After scaling, a scheduled task runs repeatedly. How do you find and remove the duplication risk? Determine whether each app replica runs its own scheduler, several CronJobs exist, a schedule overlaps, or Job retries repeat effects. Kubernetes CronJob scheduling is approximate; concurrency policy alone is not exactly-once execution. Design durable idempotency or deduplication and appropriate ownership/locking. Keep scheduled synthetic exercises local-only. See CronJob limitations↗.
ADV-34. SQS or Kafka
When would you choose a queue service versus a retained event log? Compare consumer groups, replay, ordering scope, delivery guarantees, backpressure, operations, and cost for the actual workload. Distinguish SQS queue types; neither product removes application idempotency requirements. This is elective design depth: review SQS queue types↗ and Kafka design↗, then explain the trade-off for MicroBank without pretending its messaging adapter already supports both.
Observability and incident decisions
Revisit metrics, SLIs, alerts, and postmortems.
ADV-35. Logs, metrics, and traces together
How would you design monitoring for one service and connect all three signals during diagnosis? Define a user-facing indicator, appropriate aggregations and labels, structured events, trace context, sampling, retention, and an actionable alert. Explain how you pivot from a latency metric to a request trace and supporting logs. Central platform collection and isolation are covered in PLT-13.
ADV-36. Green dashboards, unhappy users
Infrastructure and application teams disagree; dashboards and logs look healthy, but users report failure. Who do you believe? Reproduce an affected journey and collect its scope and timestamps. Check measurement coverage, freshness, sampling, missing telemetry, client/network errors, and business correctness. Treat the reports as evidence to investigate; a component health check is not an end-to-end outcome.
ADV-37. SLO alerts without constant noise
How would error budgets and multi-window burn-rate alerts change your alerting strategy? Define good/eligible events, objective and period, handling of low traffic and missing data, notification urgency, and a runbook. Explain the fast and slow windows rather than memorize thresholds. Test the rule against a labelled fixture. Use Google's SLO alerting guidance↗.
ADV-38. Thousands of alerts during an outage
How would you identify the primary failure amid an alert storm? Group by dependency, timeline, affected user journey, and shared failure domain. Correlate instead of assuming the first alert is the cause. Explain deduplication, inhibition, routing, and temporary noise control while retaining visibility into independent failures. A hypothesis needs confirming evidence.
ADV-39. Rollback fails during an outage
With a time-critical availability breach approaching, how do you decide what to do after the rollback script fails? Declare incident ownership, stop concurrent changes, preserve minimal evidence, and choose among a verified manual recovery, roll-forward, traffic shift, feature isolation, or safe degradation. Account for data compatibility and authority to act. Communicate the impact and verify recovery. Explain the uncertainty in each option rather than promising a fix within an untested deadline.
ADV-40. Incident response and your own mistake
At 2 AM the service goes down: what is your troubleshooting flow? Describe a challenging incident you resolved, including one where you introduced the failure. Distinguish immediate mitigation from causal analysis, keep a timeline, verify user recovery, and write concrete follow-up changes with owners. Use an actual experience or the local failure lab, explicitly labelled. Never present a fictional incident as personal production experience.
ADV-41. Faster releases, more incidents
What would you change when deployment speed improves but production failures increase? What recurring operational mistakes would you audit? Review change size, review/test gaps, target selection, artifact identity, schema compatibility, staged exposure, rollback readiness, observability, capacity, and ownership. Use your incident record to prioritize changes and measure both delivery and reliability; speed alone is not the outcome.
ADV-42. GitOps versus pipeline-driven deployment
How does GitOps change commit-to-deploy behavior and recovery? Separate artifact creation from a controller reconciling declared desired state. Explain drift, source credentials, promotion review, failure visibility, and emergency changes that might be reverted by reconciliation. Use the GitOps lab; multi-team controller boundaries belong in the Platform arena.
Algorithm practice with operational follow-ups
The 15-problem checkpoint is the canonical question bank. Each original problem appears once there, with its source link and corrected operational follow-up. Use these rounds to reach every supplied problem without duplicating the statements here:
| Round | Problems retained | Prepare with |
|---|---|---|
| Counting, heaps, scheduling | 347 Top K Frequent Elements; 295 Find Median from Data Stream; 621 Task Scheduler; 253 Meeting Rooms II (Premium, with an independent offline alternative) | Hashing, heaps |
| Windows and capacity | 239 Sliding Window Maximum; 76 Minimum Window Substring; 56 Merge Intervals; 1011 Capacity To Ship Packages Within D Days | Windows, binary search |
| Graph reasoning | 210 Course Schedule II; 200 Number of Islands; 684 Redundant Connection; 743 Network Delay Time | Traversal, dependency order |
| Sequence and cache reasoning | 560 Subarray Sum Equals K; 739 Daily Temperatures; 146 LRU Cache | Complexity contracts, then the checkpoint's additional patterns |
Some patterns go beyond the opening bites. Study their references before attempting them. Record your assumptions, baseline and improved solutions, complexity, edge cases, and the limits of applying each pattern to operations. Follow with the MicroBank system-design review; algorithm acceptance alone is not an application design.
Advanced project evidence
| Supplied project idea | Evidence or extension |
|---|---|
| Basic pod deployment; containerized application on Kubernetes | MicroBank kind lab, with actual request and persistence evidence. |
| Basic monitoring; Prometheus alerting; full logs/metrics/traces observability | Observability lab first; add logs and traces separately and prove correlation. |
| Centralized logging with ELK or Loki | Loki/Alloy bite; ELK is an elective alternative, not a supplied implementation. |
| Distributed tracing | OpenTelemetry bite: demonstrate context propagation across a supported application boundary. |
| HPA | Scaling bite; optional measured local exercise with a working metric source and stop condition. |
| Blue-green or canary pipeline | Canary lab; blue-green is a separate traffic-switching extension. |
| Argo CD GitOps | GitOps lab: one owner and observed reconciliation. |
| DevSecOps with SAST, DAST, and image scanning | Security lab; add a bounded local DAST exercise only against a disposable target. |
| Helm chart for an application | Elective packaging comparison with the existing Kustomize baseline; render and review first. No Helm implementation is claimed here. |
Keep periodic synthetic users and deliberate faults local-only. Run one profile at a time on the 16 GB host. For multi-region/HA design, Vault, a service mesh, disaster recovery, cost dashboards, incident automation, and the IDP capstone, continue with the Platform arena.