Skip to content
← Advanced DevOps

Practical lab guide

Lab: observe the local MicroBank journey

Export a bounded check, graph result freshness, and test a failure alert.

Documentation reviewed2026-10-01 · 4 min read · lab time varies
On this page

Start with a small measurement surface

Keep the Kubernetes baseline and both loopback API port-forwards running. Use the existing evidence/kind-case.json. This lab measures a read-only saved-case check; it does not monitor every user or continuously create new deposits.

Run Prometheus and Grafana locally using releases selected from their official installation pages. A native local Prometheus process can scrape 127.0.0.1:9108 below. If you run it in a container, its loopback is different: use an explicitly reviewed host-access mapping and test that path. Do not copy the native configuration unchanged and assume it reaches your Mac.

1. Export the result and freshness

Save scripts/study/probe_exporter.py next to the Foundations probe.py. It serves only loopback, makes at most 20 checks, and allows only one check at a time. It shuts down after the final observation window. Press Ctrl-C to stop earlier.

python
import json
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from probe import verify_case

state = {'success': 0, 'duration': 0.0, 'completed': 0.0}
lock = threading.Lock()


class Metrics(BaseHTTPRequestHandler):
    def do_GET(self):
        if self.path != '/metrics':
            self.send_error(404)
            return
        with lock:
            snapshot = state.copy()
        names = {'success': 'microbank_journey_success',
                 'duration': 'microbank_journey_duration_seconds',
                 'completed': 'microbank_journey_completed_timestamp_seconds'}
        body = ''.join(f'# TYPE {names[key]} gauge\n{names[key]} {value}\n'
                       for key, value in snapshot.items()).encode()
        self.send_response(200)
        self.send_header('Content-Type', 'text/plain; version=0.0.4')
        self.send_header('Content-Length', str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def log_message(self, *args):
        pass


def main():
    case = json.loads(Path('evidence/kind-case.json').read_text())
    server = ThreadingHTTPServer(('127.0.0.1', 9108), Metrics)
    thread = threading.Thread(target=server.serve_forever, daemon=True)
    thread.start()
    try:
        for _ in range(20):
            started = time.monotonic()
            success = 0
            try:
                verify_case(case)
                success = 1
            except (OSError, ValueError, KeyError, TypeError, RuntimeError) as error:
                print(f'Local check failed: {type(error).__name__}', flush=True)
            with lock:
                state.update(success=success, duration=time.monotonic() - started,
                             completed=time.time())
            time.sleep(30)
    finally:
        server.shutdown()
        server.server_close()


if __name__ == '__main__':
    main()

The completion timestamp updates after both passing and failing checks. A hung/unfinished check leaves it stale. Initial zero means “no completed check yet,” not a historical outage. Request timeouts remain those in the Foundations probe; a long failure can make freshness expire before the next completion.

bash
python3 scripts/study/probe_exporter.py

2. Configure a short-lived local Prometheus

Save infra/study/observability/prometheus.yaml:

yaml
global:
  scrape_interval: 15s
  evaluation_interval: 15s
rule_files:
  - rules.yaml
scrape_configs:
  - job_name: microbank-probe
    static_configs:
      - targets: [127.0.0.1:9108]

Save infra/study/observability/rules.yaml with these training rules:

yaml
groups:
  - name: microbank-study
    rules:
      - alert: MicroBankJourneyFailing
        expr: microbank_journey_success == 0
        for: 2m
        labels:
          severity: warning
          environment: local
        annotations:
          summary: "Local MicroBank journey is failing"
      - alert: MicroBankEvidenceStale
        expr: time() - microbank_journey_completed_timestamp_seconds > 120
        for: 1m
        labels:
          severity: warning
          environment: local
        annotations:
          summary: "Local journey evidence is stale"
      - alert: MicroBankProbeUnavailable
        expr: up{job="microbank-probe"} == 0 or absent(up{job="microbank-probe"})
        for: 1m
        labels:
          severity: warning
          environment: local
        annotations:
          summary: "Local probe cannot be scraped"

Use Prometheus' bundled promtool and run the server bound to loopback. Relative rule paths resolve from the configuration directory. Keep a small retention time and a private local data directory:

bash
promtool check config infra/study/observability/prometheus.yaml
promtool check rules infra/study/observability/rules.yaml
prometheus --config.file=infra/study/observability/prometheus.yaml \
  --web.listen-address=127.0.0.1:9090 --storage.tsdb.path=.local/prometheus \
  --storage.tsdb.retention.time=2h

Configure a locally running Grafana with this Prometheus data source. Bind the Grafana listener to loopback using its documented configuration. Add panels for microbank_journey_success, time() - microbank_journey_completed_timestamp_seconds, and microbank_journey_duration_seconds. Display missing data rather than substituting a green zero.

3. Exercise failure and recovery

First observe a passing result and recent timestamp. Stop only the Ledger port-forward; this deliberately breaks the observer's access path, not the Ledger service itself. Observe failed checks and the alert after its configured delay. Restore the port-forward and inspect recovery. Keep this distinction in your incident record.

To learn notification routing, start a local Alertmanager with a no-delivery training receiver, add it to Prometheus' alerting.alertmanagers configuration, and reload/restart after validation. Confirm receipt in its local UI; a receiver with no integration does not send notifications. Actual external paging is outside this lab.

Checkpoint and optional telemetry

Save the observed metrics, dashboard JSON, failure/recovery times, and what the experiment actually affected. Stop the exporter and local telemetry processes when finished. Add Alloy/Loki and OpenTelemetry using their preceding bites as separate focused exercises; a combined stack is not required on the 16 GB host. Record those integrations as complete only after verifying ingestion and correlation.

Check the meaning of each observation

Before the first completion, the exporter's zero result is an initialization value. After a completed failure, it describes a real failed read. The timestamp distinguishes them. When you stop the Ledger port-forward, the backend can remain healthy while the observer loses access; write exactly that in the record. If up becomes zero instead, inspect whether the exporter itself stopped or its scrape path broke.

The simple rules can overlap during initialization or observer failure. That is an opportunity to refine tested routing/inhibition, not evidence that several independent outages occurred. The separate rule-test fixture in the alerting bite includes an action annotation; if testing this lab's shorter rule, match the expected annotations to this exact file. Finish with one explained signal path before adding optional collectors.

Sources

Prometheus installation↗, configuration↗, Grafana installation↗, Grafana server configuration↗, Alertmanager configuration↗, and Python HTTP server↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.