Learning path / 02
DevOps foundations
Implement MicroBank locally with Python automation, AWS concepts, Terraform, Ansible, Docker, and CI/CD.
Before you begin
- Linux, networking, shell, and Git skills from System engineer foundations or equivalent experience.
- Use local exercises first; cloud exercises require an account, scoped identity, and a cleanup plan.
Your learning progress
0 of 47 available items completed
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.
The learning sequence
Explore a module to see its bites and proposed checkpoint.
01Python for automationBuild a language foundation before using Python for operational tasks.Available
- A Python project you can reproduceLearning bite
- Types, functions, and exceptionsLearning bite
- Files and configurationLearning bite
- CLI tools and subprocessesLearning bite
- HTTP APIs and testsLearning bite
- Lab: build and test a configuration-reporting CLILab guide
Produce a small tested automation tool with explicit inputs, useful errors, and documented output.
02AWS engineering foundationsEstablish cloud identity and networking before provisioning cloud resources.Available
- Accounts, IAM, and short-lived accessLearning bite
- VPC networking and DNSLearning bite
- Compute, scaling, and load balancingLearning bite
- Storage, databases, and recoveryLearning bite
- Budgets, inventory, and cleanupLearning bite
- Boto3 after IAMLearning bite
- LocalStack and local cloud dependenciesLearning bite
- Application identity: OAuth, OIDC, and tokensLearning bite
- Lab: inspect an AWS identity and design a bounded deploymentLab guide
Explain the identity used by an operation and account for the resources it creates.
03Infrastructure as code with TerraformDescribe infrastructure, inspect planned changes, and manage state deliberately.Available
- Terraform and your first local workflowLearning bite
- Providers and resource dependenciesLearning bite
- Variables and outputsLearning bite
- State and backendsLearning bite
- Modules and environment boundariesLearning bite
- Validation and change reviewLearning bite
- Lab: plan, apply, inspect, and clean up local Terraform stateLab guide
Review a plan, explain its state dependencies, and demonstrate a documented cleanup.
04Configuration management with AnsibleApply repeatable host configuration and understand idempotence.Available
- Ansible and your first playbookLearning bite
- Inventory and connectivityLearning bite
- Tasks, modules, and handlersLearning bite
- Variables and templatesLearning bite
- Roles, secrets, and check modeLearning bite
- Lab: configure a disposable host and explain the second runLab guide
Configure a disposable host and explain the second run and any changes it reports.
05Containers and local application deliveryPackage services and understand their runtime, network, and persistent data needs.Available
- Containers, runtimes, and the first lifecycleLearning bite
- Images and DockerfilesLearning bite
- Container networking and volumesLearning bite
- Compose and service healthLearning bite
- Registries and artifact identityLearning bite
- Multi-stage builds and runtime boundariesLearning bite
- Lab: run a Compose fixture and inspect MicroBank readinessLab guide
Build a MicroBank local deployment profile after checking its actual dependencies and configuration.
06CI/CD with GitHub ActionsConnect a source change to tests, an artifact, and a controlled deployment.Available
- Events, jobs, and runnersLearning bite
- Tests and quality checksLearning bite
- Credentials and artifact publicationLearning bite
- Delivery approvals and rollbackLearning bite
- Lab: turn local Python tests into a reviewable CI checkLab guide
Trace a change to a versioned artifact and its deployment evidence.
07Guided project — implement MicroBankConnect the six foundations modules through a local application deployment and recovery runbook.Available
- MicroBank 1: prepare the application projectLab guide
- MicroBank 2: build the local container profileLab guide
- MicroBank 3: provision local cloud dependenciesLab guide
- MicroBank 4: configure and start the applicationLab guide
- MicroBank 5: verify a real transaction with PythonLab guide
- MicroBank 6: build artifacts and deploy locallyLab guide
- MicroBank 7: break, recover, and close the labLab guide
Demonstrate the Accounts-to-Ledger flow, repeatable configuration, artifact identity, persistence, and scoped cleanup with actual evidence.
Automate a known baseline
Begin with Python language essentials and local automation. Learn AWS account and IAM concepts before boto3 or Terraform exercises that call cloud APIs. Terraform and Ansible then provide a useful division between provisioning resources and configuring hosts.
GitHub Actions is the primary CI/CD implementation for this path. Jenkins and GitLab CI are later comparison electives. GitOps reconciliation and Argo CD follow Kubernetes in Advanced DevOps. See the GitHub Actions workflow model↗.
The six foundations modules contain 34 learning bites and six preparatory labs. A seventh module adds seven ordered MicroBank implementation labs, giving this path 47 available items. Learn each tool with a small exercise, then use the same concepts in the application project.
Learn each tool before connecting the application
Each module begins with its purpose and vocabulary, then a small worked example, practice, and questions with reasoning guides. Python starts with a reproducible project and the language itself. Terraform first completes a local plan/apply workflow. Ansible first changes one local file. Docker first explains the image, process, runtime, and lifecycle. Their capstone use should therefore feel like connecting familiar pieces.
AWS practice includes account-free design work and an offline SDK response fixture. Real AWS reads require your assigned sandbox; no paid stack is needed for these foundations. Read the LocalStack concept lesson in the cloud module, then return to its runtime checks after Docker/Compose. The seven implementation guides supply the actual local application profile.
Keep expected output separate from observations. Record a check only after running it, and explain what it establishes. A unit test, rendered configuration, healthy process, successful transaction, and tested recovery are different milestones. The lessons are reviewed; they are not a claim that your selected environment has already passed the labs.
This path covers foundations rather than every source chapter. Python packaging/async clients, advanced Ansible plugins and execution strategies, manual namespace/container construction, Podman, managed AWS container services, deep CloudWatch/KMS operations, and alternate CI platforms remain optional depth or later-path work.
Implement MicroBank here
System engineer foundations uses your personal learning site. DevOps foundations implements MicroBank as its guided application project. Begin with the project setup and scope, then follow the seven implementation labs after the concept modules. The guides include the configuration and code you will add to the application.
The local target is Accounts → SNS/SQS in LocalStack → Ledger, backed by PostgreSQL. The acceptance check creates synthetic data, submits a deposit, repeats its idempotency key, and verifies a matching Ledger entry and balance. The optional frontend retains its actual Auth0 login. Missing settlement consumption, real notifications, and frontend balance integration are recorded as application backlog items.
| Stage | What you implement | Evidence |
|---|---|---|
| Prepare | Pinned source, application inventory, explicit Accounts publisher lifecycle | Focused branch and reviewed diff |
| Package | Standalone Compose profile, service images, private databases and named volumes | Healthy local dependencies and resolved image identities |
| Provision | Local topics, queues, policies, subscriptions and an S3 practice bucket | Reviewed Terraform plan and no-change follow-up |
| Configure | Runtime configuration generated from infrastructure outputs | A repeatable Ansible run and service readiness |
| Verify | Bounded Python transaction probe and failure tests | Real Ledger result, repeated-key check and unit test results |
| Deliver | Manual GitHub Actions image builds and selected local deployment | Source SHA, CI result, image IDs and release note |
| Operate | Consumer interruption, restart verification and scoped cleanup | Recovery runbook and observed persistence |
The implementation uses MicroBank's inspected source↗, with original lab configuration written for this path. The LocalStack concept bite explains local AWS emulation before the implementation steps. Choose and record a compatible LocalStack image and authentication setup; emulator behavior does not establish real AWS behavior.
On a 16 GB host, run one track at a time and build sequentially. Synthetic transaction checks stay local. Hosted CI runs unit checks and image builds only. Independent VM deployment, real-cloud adapters, Kubernetes, GitOps, advanced observability and progressive delivery continue in the later learning paths and Build, Break & Operate.
Publish the commands, versions, observations, failures and recovery evidence on your personal learning site. These are implementation guides with explicit acceptance checks; reviewed content does not mean a learner's chosen environment has already passed them.
Relation to certification study
This is an engineering learning path. The Terraform Associate study workspace retains its own exam objectives, practice questions, revision material, and saved progress. Use that workspace for exam-specific recall and practice; the engineering labs keep their own progress.
Additional roadmap coverage
Week 1 of the AI Platform Engineering Handbook↗ informs original additions on application identity and build/runtime boundaries, plus stronger CI evidence, cache/artifact, and image-identity exercises. Week 4's Terraform material informs the module/state handoff to Platform Engineering. Application identity also addresses the supplied Linux → networking → auth → containers roadmap: cloud IAM, browser login, API token validation, and account authorization are different concerns.
The source is supplementary reading, not a copied course or evidence of a working deployment. Its AI/ML examples and assumed productivity targets are outside this path. Here, you move from a personal learning site to Python automation and the MicroBank application.
Source and adaptation
The topic structure draws on the collaborative Compute Central curriculum↗, adapted with permission. The explanations and MicroBank exercises are written for this path, with primary documentation linked in each lesson. Python starts here after System engineer foundations. Lab completion depends on running and checking the exercises in your own environment.
Interview prep arena
Use your documented MicroBank revision for concrete examples. Attempt each question aloud, then retain the evidence, the alternative explanation you ruled out, and a safe verification step. These scenarios are for practice; they do not describe observed MicroBank incidents or predict an employer's questions. Related questions are grouped together so you can work through a topic and its follow-ups.
Choose a round: Containers · CI/CD · Terraform · AWS · Project evidence.
Containers and artifacts
Revisit Dockerfiles, runtime boundaries, and artifact identity.
DEV-01. What docker run does
Trace docker run from the client request to the running process. Explain image resolution and pull policy, container creation, writable filesystem, mounts, networking, isolation, configured command, and exit behavior. On your Mac, identify the Linux environment provided by the container engine; Linux containers do not acquire a separate kernel each.
DEV-02. CMD and ENTRYPOINT
How do CMD and ENTRYPOINT combine, and what can a caller override? Contrast exec and shell forms, argument handling, and signal delivery to the main process. Use a tiny disposable image to show the effective command instead of relying only on definitions.
DEV-03. Layers and forty image builds
How does layer caching work, and how would you improve a build of 40 images taking 18 minutes? Inspect cache misses, Dockerfile ordering, context size, shared dependencies, BuildKit cache mounts, and external cache reuse. Rebuild selectively using real dependencies, including shared library changes. Measure bounded parallelism; more concurrent builds can contend for the same CPU and disk. Revisit Docker cache guidance↗.
DEV-04. Works locally, fails after deployment
How would you investigate an image that works locally but fails in production or only after promotion from staging? Compare the actual digest and architecture, command, UID, configuration, secret references, DNS, dependency versions, mounts, limits, and network access. Distinguish a build difference from an environment contract difference. Preserve failure evidence before changing several settings at once.
DEV-05. Mutable and immutable infrastructure
When would you update a running host and when would you replace it from a versioned image? Discuss drift, patching, rollout time, data persistence, and recovery for each approach. Explain why replacing an application image does not make its database stateless, and connect the decision to Ansible configuration and Terraform ownership.
Delivery and pipeline investigation
Revisit jobs and runners, quality checks, credentials, and delivery recovery.
DEV-06. Commit to a verified deployment
Walk through CI/CD end to end, including its internal handoffs. Trace trigger and revision → checkout → tests → artifact build and publication → approval → target selection → deployment → user-facing verification. Identify credentials, outputs, failure handling, and the exact digest at each boundary. Explain which portions your current lab actually executes.
DEV-07. Parallel jobs and dependencies
How do pipeline jobs run concurrently while respecting prerequisites? Draw the job DAG; distinguish independent jobs, dependency results, artifact transfer, conditional execution, cancellation, and environment concurrency controls. A dependency edge orders work but does not make deployment idempotent. Compare your answer with GitHub Actions jobs↗.
DEV-08. A pipeline takes too long
How would you diagnose a 25–40+ minute pipeline and assess a request to get below 10 minutes without adding hardware? Separate queue time, critical-path work, test duration, image builds, downloads, and deployment waits. Consider cache validity, test sharding, dependency-aware selection, and build-once promotion. Retain required checks; report measured improvement and the remaining lower bound rather than guaranteeing a target. Use DEV-03 for image-specific diagnosis.
DEV-09. Green pipeline, wrong target or failed delivery
A pipeline succeeds, but deployment fails or reaches a different environment. What do you verify? Check whether the delivery job ran, skipped, masked an error, or returned before rollout completion. Trace account, region, cluster context, namespace, overlay, branch, environment variables, and approvals against the intended target. When the target is right but traffic still reaches an old version, continue with ADV-09.
DEV-10. Secure pipeline credentials and artifacts
How would you secure CI/CD and its secrets? Cover short-lived scoped credentials, protected environments, untrusted pull requests, runner isolation, action/dependency pinning, cache trust, artifact provenance, and log handling. Explain how a deployment identity differs from a test identity and how you revoke access when a runner is compromised.
DEV-11. Rollback after failed delivery
How would you recover a failed deployment safely? Identify a known-good artifact and configuration, verify database compatibility, coordinate the deployment owner, and check the user journey after restoration. State when roll-forward or temporarily disabling a feature is safer than rollback. The escalation variant—rollback itself fails during an outage—is retained in ADV-39.
DEV-12. A secret was committed and cloned
What are your next actions when a credential has escaped into Git history? Revoke or rotate it promptly, assess access and blast radius, update dependent systems, and preserve relevant audit evidence. Coordinate history cleanup and downstream copies; deleting a commit cannot retract a cloned secret. Add prevention checks and verify the old credential no longer works. See GitHub's remediation guidance↗.
Terraform operating basics
Revisit dependencies, state, and change review.
DEV-13. Dependency graph and hidden ordering
How does Terraform construct and walk its dependency graph? When is depends_on needed? Contrast dependencies inferred from references with hidden behavioral dependencies. Explain why an explicit dependency may be necessary and why a broad module-level dependency can make planning more conservative. Do not assume Terraform can infer every application readiness requirement. Graph internals↗ and depends_on↗ provide the implementation boundaries.
DEV-14. State locking and a lost lock
Why is locking needed, which operations use it, and what would you do if a lock disappeared mid-apply? Identify the actual backend and its failure semantics; there is no universal lease-loss behavior to invent. Stop new writers, establish whether the original run is still active, preserve its diagnostics and state versions, then reconcile remote objects with recorded state. Never force-unlock an active owner's lock. Locking↗ prevents competing state writers where supported; it is not an infrastructure transaction.
DEV-15. count, for_each, and identity changes
How do indexed instances differ from keyed instances, and why can switching between them propose destruction? Explain address identity and list reordering. Show how to review the old-to-new address mapping before a refactor; do not assume similar arguments preserve identity. Continue with PLT-18 for a reviewed migration.
DEV-16. Moving state without recreating infrastructure
How would you migrate a backend or separate state ownership without recreating resources? Distinguish backend migration from moving addresses between configurations. Coordinate writers, retain protected backups, inspect lineage and resource addresses, and review a no-unintended-change plan afterward. Follow the relevant initialization↗ and state refactoring↗ procedures; manually copying files is not the whole migration.
DEV-17. Drift, missing resources, and a live service
How do you detect drift, including when someone reports a manual change but Terraform shows no difference? How would you reconcile it while the service is critical? Verify account, region, workspace, state address, refresh settings, ignored attributes, and whether the changed object is managed at all. Compare configuration, state, and provider-observed reality. Drift is not automatically an apply error; inspect the failing operation and provider message. A manually deleted managed object may be planned for recreation if it remains required; inspect the plan. Decide whether to adopt or revert the change, reviewing replacement and availability risks. Refresh-only updates records, not configuration or remote infrastructure. Use the plan reference↗.
DEV-18. A partially completed apply
What happens when apply fails halfway, and how do you recover? Completed remote changes are not automatically rolled back. Preserve diagnostics and state, inspect actual resources and the failure, correct the cause, and review a fresh plan before continuing. A state-write failure needs special reconciliation; do not retry destructive actions blindly. See apply recovery guidance↗.
Cloud request and access boundaries
Revisit IAM, AWS networking, and the read-only lab.
DEV-19. IAM roles and policies
How does an IAM role differ from a policy? Explain an assumable identity, its trust relationship, temporary credentials, and permission evaluation. Give a scoped workload example; possession of a role name is not permission to assume it.
DEV-20. VPC, subnets, and gateways
Explain a VPC, subnet routes, internet gateway, and NAT gateway using a request and its return path. Identify public-address requirements, security groups, network ACLs, and DNS. Then diagnose an EC2 instance whose health checks pass but whose application is unreachable, or two provisioned services that cannot communicate. Include the application's listening address and port rather than stopping at network configuration.
DEV-21. AWS load balancer routing
How does an AWS load balancer select a target? First identify ALB versus NLB and the relevant protocol. Explain listeners, applicable routing rules, target groups, health checks, and connection handling for that type. Propose evidence to distinguish a load-balancer response from a target response. See Elastic Load Balancing concepts↗.
DEV-22. A CloudFront request
What happens when a client requests a CloudFront URL? Follow DNS, TLS, cache behavior and cache key, hit versus miss, origin request, and response caching. Explain how headers, query strings, cookies, TTLs, and invalidation can affect the observed version. CloudFront is elective depth here; start with its request flow↗.
DevOps project evidence
Use one local MicroBank profile at a time. Choose a project that helps you practice a skill and collect your own results. These ideas are optional extensions, not additional verified implementations.
| Project idea | Existing route or explicit extension |
|---|---|
| Dockerize a basic application | Compose implementation: retain the image digest, startup command, and health evidence. |
| Simple GitHub Actions or Jenkins pipeline | CI evidence lab uses Actions. Jenkins is an optional reimplementation, not a supplied lab. |
| Full build → test → deploy pipeline | MicroBank delivery lab, with its local-deployment boundary. |
| Terraform for EC2, VPC, and S3 | Read-only AWS design plus the local Terraform implementation. Actual EC2/VPC/S3 provisioning is an optional costed extension. |
| Static site on S3 and CloudFront | Elective continuation of the personal site: review origin access, TLS, cache behavior, cost inventory, and teardown before any live deployment. |
For reusable modules, multi-team state, provider upgrades, large plans, secrets, and state recovery, use the Platform Terraform round. For algorithm interviews after Python, use the Advanced DSA round.