Learning path / 04
Platform engineering
Build a supported MicroBank platform with environment contracts, GitOps, Kyverno, target adapters, and a developer portal.
Before you begin
- A documented MicroBank application profile with actual local build, transaction, and recovery evidence from the earlier paths.
- Understand Kubernetes, Argo CD, artifact identity, configuration, and the difference between application and infrastructure checks.
- Resolve the manifest/controller handoff before applying platform examples; use one local profile at a time on the 16 GB host.
Your learning progress
0 of 33 available items completed
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.
The learning sequence
Explore a module to see its bites and proposed checkpoint.
01Platform users and service ownershipDefine the supported developer journey before designing a portal.Available
- Users and recurring delivery tasksLearning bite
- Service ownership and contractsLearning bite
- Documentation and support boundariesLearning bite
- Platform feedback and adoptionLearning bite
- Lab: define the MicroBank platform contractLab guide
Describe a developer task, its inputs, support owner, and an observable successful result.
02Environment configuration and promotionPromote the same artifact through explicit environment differences.Available
- Kustomize bases and overlaysLearning bite
- Dev, QA, preprod, and production gatesLearning bite
- Configuration and secret boundariesLearning bite
- Promotion and rollback recordsLearning bite
- Lab: render MicroBank environment contractsLab guide
Trace one immutable artifact through reviewed environment changes.
03GitOps at platform scopeProvide repeatable delivery with clear access and reconciliation boundaries.Available
- Argo CD application organisationLearning bite
- Repository and environment boundariesLearning bite
- Controller access and recoveryLearning bite
- Reusable deployment conventionsLearning bite
- Reusable delivery workflowsLearning bite
- Lab: hand off one MicroBank GitOps ownerLab guide
Demonstrate an auditable environment change and document the recovery procedure.
04Policy as code with KyvernoMake supported workload rules testable and explainable.Available
- Admission and background checksLearning bite
- Audit before enforcementLearning bite
- Policy tests and exceptionsLearning bite
- Developer feedbackLearning bite
- Lab: test and introduce a MicroBank owner policyLab guide
Show an accepted and rejected workload and explain the relevant policy result.
05Pluggable deployment targetsKeep application intent consistent while making provider differences explicit.Available
- Local deployment contractLearning bite
- EKS or GKE integrationLearning bite
- Identity, networking, storage, and ingressLearning bite
- Cost inventory and teardownLearning bite
- Terraform modules and state boundariesLearning bite
- Lab: review a MicroBank cloud adapterLab guide
Review one target adapter; distinguish design, implementation, and live verification, including cost and cleanup evidence.
06Internal developer platform capstoneConnect one developer portal to delivery automation that already works.Available
- Backstage or Port evaluationLearning bite
- Service catalog and ownershipLearning bite
- Templates and self-service actionsLearning bite
- Pipeline, policy, and runtime statusLearning bite
- Approvals, audit, and bounded automationLearning bite
- Lab: request a MicroBank change through an IDPLab guide
Request a supported change through the portal, follow its execution, and demonstrate a documented failure and recovery path.
Turn the working project into a supported platform
Platform engineering has 27 learning bites and six practical labs across six modules. Continue MicroBank from DevOps Foundations and Advanced DevOps. You’ll turn a delivery task you have already demonstrated into something other developers can request, review, operate, and recover through a supported workflow.
Bites introduce the terms, work through an example, and give you a small decision or exercise before the larger lab. Diagrams explain delivery, overlays and recovery; reasoning guides let you compare your answer. Read the concept bites in module order, then use the module lab to combine them. Each lesson is marked Reviewed. To record a lab as executed, run its documented procedure and retain the results. Include the source repository, application profile, installed tool versions, and observed outcomes in that record.
Learn the task, then make the handoff
Start with the four bites on developer tasks, ownership, support and feedback. They include worked decisions and a visual delivery journey. Then complete the platform contract lab. Earlier guides use MicroBank revision f75da68 and learner-created configuration. A different working branch may add a settlement worker or frontend behavior; those changes require a corresponding deployment contract. Do not assume all versions of MicroBank have the same services or endpoints.
Reconcile the last working manifests before adding another controller. Preserve hardening changes, choose one canonical base, and establish one owner per workload. If you completed the temporary Rollout exercise, restore and verify the Deployment baseline before starting this path’s platform handoff.
The practical progression
| Module | Learning bites | Lab and evidence |
|---|---|---|
| Platform users and ownership | 4 | Define the contract: supported task, service inventory, ownership, manual baseline |
| Environment configuration | 4 | Render five overlays: one base, local runtime, four review-only stage declarations |
| GitOps at platform scope | 5 | Hand off the local owner: restricted source/destination, reviewed sync, recovery evidence |
| Kyverno policy | 4 | Test an ownership rule: offline positive/negative cases, optional local Audit/Deny verification |
| Deployment targets | 5 | Review a cloud adapter: identity, dependencies, state, costs, acceptance, teardown |
| IDP capstone | 5 | Request a change through a portal: real catalog, constrained request, PR, checks, sync, verified result |
The cloud lab guides you through a design and provisioning review, with an optional later live session. Building and verifying an EKS/GKE provisioning implementation is separate work. The portal lab gives a concrete Backstage request template and a Port adaptation contract. Its first annotation-only request tests the integration; a complete artifact release requires the subsequent release-selection implementation and evidence.
Keep the resource and test boundaries
On a 16 GB RAM / 512 GB storage Mac, run one local profile at a time. Render dev, QA, preprod, and production declarations without running four full clusters. Pause optional telemetry or the portal when testing controllers; measure actual memory, disk, and startup behavior. A single-node kind cluster and a separate two-node kubeadm lab are operating exercises, not HA platforms.
Recurring synthetic users and deliberate faults remain local-only. They are excluded from the platform application base, cloud targets, hosted CI, and portal templates. Static checks can run in CI. No cloud cost, performance improvement, or successful rollout is inferred from a generated file.
Cloud portability needs application decisions
MicroBank uses AWS messaging semantics. EKS and GKE can schedule Kubernetes workloads, but GKE does not translate SNS/SQS into Pub/Sub. Choose one target and document how identity, databases, queues, storage, private access, and observability work there. The inspected Ledger code's credential construction needs review before adopting workload identity. Backend authorization and ledger consistency remain application concerns, even when policy checks pass.
Original lessons with explicit references
The curriculum uses newly written explanations and MicroBank-specific exercises. Week 1↗ of the AI Platform Engineering Handbook informs delivery, artifact, security, Kubernetes, and observability prerequisites. Week 4↗ informs platform product thinking, Terraform interfaces, GitOps promotion, and self-service. We adapt applicable concepts rather than copy its chapters, diagrams, code listings, AI workloads, or numerical outcome claims. Its MIT license↗ and author attribution are retained in the source map.
The collaborative Compute Central curriculum↗ also informs this path, adapted with permission. Every lesson links primary documentation. The Ship Boring capstone records the wider project direction; this course does not claim it is already deployed.
Interview prep arena
Use the working MicroBank project to discuss platform decisions, ownership, and failure recovery. These questions build on the earlier operating skills and introduce systems beyond the course labs. Label service mesh, operator, multi-tenant EKS, VPN, etcd tuning, and multi-region work as design extensions unless you have independent execution evidence.
Choose a round: Platform boundaries · Architecture · Cost · Terraform · Project evidence.
Platform ownership and security
Revisit Argo CD organisation, environment contracts, secret boundaries, and policy tests.
PLT-01. Independent team releases with GitOps
How would you structure GitOps for several teams that release independently? Compare ApplicationSets, application grouping, and an app-of-apps pattern against ownership and trust boundaries. Define repository permissions, AppProjects, permitted destinations, promotion review, controller credentials, and recovery. Generated applications do not create authorization boundaries by themselves. Follow Argo CD security guidance↗ and its ApplicationSet documentation↗.
PLT-02. Repeated deployments produce inconsistent infrastructure
How would you investigate and prevent different deployments from disagreeing about desired state? Identify each writer: Terraform, Ansible, pipelines, GitOps controllers, and manual changes. Compare artifact versions, configuration inputs, target identity, state ownership, and concurrent runs. Define one authoritative owner per resource, explicit handoffs, reviewed environment differences, and verifiable promotion. State locking cannot coordinate unrelated controllers that own the same object.
PLT-03. Secret delivery for many services
How would you serve 50+ workloads without distributing broad Vault access? Compare per-service Kubernetes authentication and narrow Vault roles, Agent injection, dynamic credentials/lease renewal, and an External Secrets Operator integration. Explain rotation, revocation, unavailable-secret behavior, and the operator's authority. Injector delivery does not remove authentication requirements; synchronization to Kubernetes Secrets changes where values persist. See Vault injection↗, Kubernetes authentication↗, and ESO's Vault provider↗. AWS Secrets Manager is an alternative to evaluate against the same contract.
PLT-04. Multi-tenant EKS isolation
How would you isolate tenants on a shared EKS cluster? Establish the tenant threat model, then examine RBAC, namespaces, network isolation, workload identity, admission, quotas, node isolation, and tenant-scoped telemetry access. Include noisy-neighbor and controller-access concerns. Decide when the trust requirement warrants separate clusters or accounts; namespaces alone do not provide every isolation property. Use Kubernetes multi-tenancy guidance↗.
PLT-05. External database reached through a VPN
How would you design secure, highly available private database access from Kubernetes? Compare a network-managed tunnel with an in-cluster tunnel appliance, stating routing, DNS, source addressing, MTU, encryption, authentication, failover, and operational ownership. Review connection recovery and database-side restrictions. Do not assume one VPN sidecar provides HA or that connectivity alone authorizes data access. Document the selected provider and failure domains before producing manifests.
PLT-06. Service mesh versus Ingress
How do their responsibilities differ, and when would a mesh earn its operational cost? Distinguish traffic entering the application boundary from service-to-service identity, telemetry, and traffic policy. Identify overlap with gateways and proxies, and compare the simpler design that meets the requirements. This is elective depth; consult Istio concepts↗ or Linkerd architecture↗ for the implementation you select.
PLT-07. A sidecar costs more than the application
How would you investigate and reduce Envoy or another mesh proxy's CPU/memory use? Measure per-proxy traffic, connections, configuration size, telemetry overhead, and resource settings. Compare a representative baseline, consider configuration scoping, and validate latency/security after any change. Removing a sidecar can remove relied-on protections. Use the selected mesh's performance guidance↗ rather than promise a fixed reduction.
PLT-08. Design an operator
How would you design a CRD and controller for a complex application lifecycle? Define desired state, validation, status conditions, reconciliation, retries, idempotency, resource ownership, finalizers, and upgrade compatibility. Separate asynchronous progress from a successful API response. Test interrupted operations and cleanup. Review the operator pattern↗; building an operator is an extension, not a prerequisite for the current portal lab.
PLT-09. Degrading etcd performance
How would you investigate etcd latency and design its availability? Examine storage fsync latency, network RTT, leader changes, database size, request pressure, quorum, and the supported topology. Separate measurement from maintenance such as compaction/defragmentation, and rehearse backup/restore before intervention. Respect managed-control-plane boundaries. Read etcd performance↗ and maintenance↗; changing member count or tuning blindly is not a diagnosis.
PLT-10. Trusted registries at admission
How would you enforce approved image sources, and verify that enforcement works? Define an exact registry/repository rule, applicable workload kinds and container fields, exception ownership, and policy rollout. Test allowed and rejected images, including misleading prefix matches. Registry location is separate from signature/provenance verification. Extend the Kyverno lab using the installed version's supported policy API.
Platform architecture and recovery
Revisit supported tasks, workflow reuse, controller recovery, and delivery status.
PLT-11. Highly available CI/CD
How would you design a CI/CD platform that can survive component failure? Map the control plane, queues, runners, artifact store, secrets, approvals, and deployment ownership. Consider retries, duplicate jobs, runner isolation, persisted state, recovery objectives, and regional dependencies. Distinguish durable delivery records from disposable runner caches; use a failure exercise to test the chosen boundary.
PLT-12. Highly available infrastructure
How would you design HA across nodes, zones, or regions, and choose autoscaling boundaries? Begin with user impact, failure domains, data replication, traffic routing, dependencies, and cost. Explain where horizontal and vertical scaling help and where they do not supply redundancy. Multi-region is an elective design exercise; the local kind and two-node kubeadm profiles do not demonstrate production HA.
PLT-13. Central observability platform
How would you provide shared logs, metrics, and traces for many services or tenants? Define collection, buffering, correlation, retention, cardinality, access controls, cost limits, and failure behavior. Include telemetry freshness and alerts when the monitoring path itself fails. Plan the evidence needed to investigate one request across services; installing three dashboards is not a complete observability outcome.
PLT-14. Disaster recovery
How would you design and test backup, restoration, and failover? Agree on RPO/RTO for data and service recovery, inventory dependencies and credentials, protect backups, and test an isolated restore followed by an application-level consistency check. Include traffic cutover, authority during partitions, failback, and operator access. HA is not a backup strategy. Use recovery-objective guidance↗; label untested times as targets.
Cost investigation
Revisit cost inventory and teardown and the cloud-adapter review.
PLT-15. AWS spend triples without a deployment
What is your step-by-step response to a sudden bill increase? Check reporting period and data freshness, then break down service, account, region, usage type, and resource/tag attribution. Compare request/traffic growth, data transfer/NAT, storage and logs, autoscaling, orphaned resources, pricing/discount changes, and unauthorized use. Correlate audit and usage evidence before stopping critical resources. Start with Cost Explorer↗.
PLT-16. Cost reduction without harming availability
How would you reduce cloud spend while preserving performance and availability requirements? Establish demand, service objectives, headroom, recovery requirements, and the largest attributable costs. Evaluate sizing, scheduling nonproduction work, storage/log retention, transfer paths, and purchasing options against actual usage. Test a bounded change and report its impact; a cheaper instance or fewer replicas is not automatically an optimization.
Terraform at team scale
Revisit Terraform platform contracts. The Foundations arena owns the canonical questions on DAGs, locking/lost locks, count/for_each, state migration, drift/manual deletion, and partial applies; those questions are not repeated here.
PLT-17. A 200 MB state and a twelve-minute plan
How would you locate the cost of a large plan and improve it? Measure state transfer/serialization, refresh/provider API latency, throttling, graph size, data sources, and runner limits before choosing a remedy. Separate configurations/state by ownership and lifecycle with explicit contracts when justified. Creating child modules alone does not split state. Review migration and cross-state dependencies; a remote backend or higher parallelism is not a guaranteed speedup.
PLT-18. Refactoring resource addresses
How do you refactor modules, rename resources, or move from count to for_each without unintended replacement? Map old and new addresses and choose a supported moved declaration or a coordinated state operation. Check provider/resource compatibility and inspect a plan for unexpected destruction. Distinguish address moves inside a configuration from transfers between state owners. Use HashiCorp's refactoring guide↗.
PLT-19. No-change plan followed by an unexpected modification
How would you investigate this report without assuming Terraform ignored its plan? Determine whether the user applied the exact saved plan or ran a new automatic plan. Compare revision, inputs, workspace, credentials, provider selection, flags, external changes, and timestamps. Separate state/refresh changes from remote mutations and inspect cloud audit evidence. A verified no-op saved plan is not a normal instruction to mutate managed resources; investigate other writers or defects and describe only the evidence you have. See plan↗ and apply↗.
PLT-20. Terraform secrets and persistence
How do you avoid unnecessary secret persistence, and what if the provider needs to store a value? sensitive redacts display but does not omit values from state. Review supported ephemeral values and write-only arguments, their Terraform/provider requirements, and permitted reference contexts. Otherwise protect state and plans, or move secret delivery outside Terraform when appropriate. A secret-manager lookup alone does not guarantee omission. Use sensitive-data guidance↗.
PLT-21. State across teams and environments
How would you organize safe multi-team state ownership? Define separate roots/backends or state keys, identities, permissions, locking, encryption, versioned recovery, audit, and serialized runs for each ownership boundary. Limit the data shared between configurations. Explain how environment selection is validated before apply and how simultaneous changes are coordinated.
PLT-22. Modules reference the same resource
When is sharing a resource reference harmless, and when does it create competing ownership? Multiple consumers can use a resource ID or output. Multiple managed addresses claiming the same remote object create a different problem. Name one owner, pass explicit interfaces, and plan a controlled ownership transfer if needed; do not import the same object into several states.
PLT-23. Reusable modules without tight coupling
How would you structure modules for multiple environments? Separate reusable resource behavior from root-level wiring, credentials, backend choice, and environment inputs. Keep outputs purposeful, document dependencies and version constraints, and test compatibility. Explain trade-offs of module granularity and avoid assuming that one module per team necessarily matches a state boundary.
PLT-24. Workspaces and organizational boundaries
How do CLI workspaces work, and where can they be misused? Explain separate state instances for one configuration/backend and the risk of selecting the wrong instance. They are not inherently dangerous, but they do not provide strong authorization or system-decomposition boundaries. Distinguish CLI workspaces from HCP Terraform workspaces. See workspace guidance↗.
PLT-25. Provider version mismatch
How can inconsistent provider versions break a workflow, and how would you prevent it? Review constraints, selected versions and checksums in the dependency lock file, supported Terraform versions, provider schema changes, and upgrade testing. Commit and review lock-file changes, including appropriate platform checksums. Do not treat a provider downgrade as a universal state-recovery method. Read dependency locking↗.
PLT-26. State corruption and recovery evidence
How would you respond to damaged state, and how would you explain a real incident if you have experienced one? Stop competing writers, protect the current evidence, inspect backend versions and object identity, and reconcile a validated recovery snapshot against real infrastructure. A backup can be stale. Use supported import/state procedures under a reviewed plan; avoid blind state pushes or destructive retries. If you have no incident, present a disposable recovery rehearsal honestly. See state backup and recovery↗.
Platform project evidence
Use this project list to choose a next step. Review its prerequisites and resource needs before starting one extension; a linked design bite will help you plan the implementation.
| Project idea | Evidence or explicit extension |
|---|---|
| End-to-end microservices CI/CD platform | Join the working earlier labs to reusable workflows, environment gates, and a verified delivery record. |
| Multi-environment infrastructure with Terraform modules | Terraform contracts and environment lab; retain separate ownership and rendered differences. |
| Secure secrets with Vault or AWS Secrets Manager | PLT-03 design extension: demonstrate scoped access, denied access, rotation, and failure behavior for one chosen integration. |
| Kubernetes autoscaling and HA | PLT-12 design extension; managed nodes, extra control-plane capacity, and HA storage require a separate implementation and budget. |
| Multi-region deployment | PLT-12/14 design extension with replication, cutover, failback, and explicit RPO/RTO. Do not run multiple cloud regions just to answer a question. |
| Istio or Linkerd service mesh | PLT-06/07 elective: measure one justified capability and its resource overhead. |
| Full observability platform and distributed logging/tracing | PLT-13 architecture extends the Advanced local telemetry evidence with tenant access, retention, and collector failure behavior. |
| Disaster recovery backup and failover | PLT-14: a bounded local restore rehearsal first, then a reviewed target-specific design. |
| Cloud cost dashboard | PLT-15/16: use a sanitized billing export or labelled fixture with attribution and data freshness; do not invent savings. |
| Incident-response automation | Bounded automation: begin with read-only evidence collection, scoped authority, audit, approvals where required, and a stop path. |
| Internal Developer Platform | Backstage or Port capstone, connected to a delivery operation that already works. |
For an algorithm follow-up to dependency orchestration, capacity planning, or caching, return to the single Advanced question bank. Keep assumptions, complexity, and real-system limits visible. All scheduled synthetic users and deliberate fault exercises remain local-only.