Practical lab guide
Lab: observe the local MicroBank journey
Export a bounded check, graph result freshness, and test a failure alert.
On this page
Start with a small measurement surface
Keep the Kubernetes baseline and both loopback API port-forwards running. Use the existing evidence/kind-case.json. This lab measures a read-only saved-case check; it does not monitor every user or continuously create new deposits.
Run Prometheus and Grafana locally using releases selected from their official installation pages. A native local Prometheus process can scrape 127.0.0.1:9108 below. If you run it in a container, its loopback is different: use an explicitly reviewed host-access mapping and test that path. Do not copy the native configuration unchanged and assume it reaches your Mac.
1. Export the result and freshness
Save scripts/study/probe_exporter.py next to the Foundations probe.py. It serves only loopback, makes at most 20 checks, and allows only one check at a time. It shuts down after the final observation window. Press Ctrl-C to stop earlier.
import json
import threading
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path
from probe import verify_case
state = {'success': 0, 'duration': 0.0, 'completed': 0.0}
lock = threading.Lock()
class Metrics(BaseHTTPRequestHandler):
def do_GET(self):
if self.path != '/metrics':
self.send_error(404)
return
with lock:
snapshot = state.copy()
names = {'success': 'microbank_journey_success',
'duration': 'microbank_journey_duration_seconds',
'completed': 'microbank_journey_completed_timestamp_seconds'}
body = ''.join(f'# TYPE {names[key]} gauge\n{names[key]} {value}\n'
for key, value in snapshot.items()).encode()
self.send_response(200)
self.send_header('Content-Type', 'text/plain; version=0.0.4')
self.send_header('Content-Length', str(len(body)))
self.end_headers()
self.wfile.write(body)
def log_message(self, *args):
pass
def main():
case = json.loads(Path('evidence/kind-case.json').read_text())
server = ThreadingHTTPServer(('127.0.0.1', 9108), Metrics)
thread = threading.Thread(target=server.serve_forever, daemon=True)
thread.start()
try:
for _ in range(20):
started = time.monotonic()
success = 0
try:
verify_case(case)
success = 1
except (OSError, ValueError, KeyError, TypeError, RuntimeError) as error:
print(f'Local check failed: {type(error).__name__}', flush=True)
with lock:
state.update(success=success, duration=time.monotonic() - started,
completed=time.time())
time.sleep(30)
finally:
server.shutdown()
server.server_close()
if __name__ == '__main__':
main()
The completion timestamp updates after both passing and failing checks. A hung/unfinished check leaves it stale. Initial zero means “no completed check yet,” not a historical outage. Request timeouts remain those in the Foundations probe; a long failure can make freshness expire before the next completion.
python3 scripts/study/probe_exporter.py
2. Configure a short-lived local Prometheus
Save infra/study/observability/prometheus.yaml:
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- rules.yaml
scrape_configs:
- job_name: microbank-probe
static_configs:
- targets: [127.0.0.1:9108]
Save infra/study/observability/rules.yaml with these training rules:
groups:
- name: microbank-study
rules:
- alert: MicroBankJourneyFailing
expr: microbank_journey_success == 0
for: 2m
labels:
severity: warning
environment: local
annotations:
summary: "Local MicroBank journey is failing"
- alert: MicroBankEvidenceStale
expr: time() - microbank_journey_completed_timestamp_seconds > 120
for: 1m
labels:
severity: warning
environment: local
annotations:
summary: "Local journey evidence is stale"
- alert: MicroBankProbeUnavailable
expr: up{job="microbank-probe"} == 0 or absent(up{job="microbank-probe"})
for: 1m
labels:
severity: warning
environment: local
annotations:
summary: "Local probe cannot be scraped"
Use Prometheus' bundled promtool and run the server bound to loopback. Relative rule paths resolve from the configuration directory. Keep a small retention time and a private local data directory:
promtool check config infra/study/observability/prometheus.yaml
promtool check rules infra/study/observability/rules.yaml
prometheus --config.file=infra/study/observability/prometheus.yaml \
--web.listen-address=127.0.0.1:9090 --storage.tsdb.path=.local/prometheus \
--storage.tsdb.retention.time=2h
Configure a locally running Grafana with this Prometheus data source. Bind the Grafana listener to loopback using its documented configuration. Add panels for microbank_journey_success, time() - microbank_journey_completed_timestamp_seconds, and microbank_journey_duration_seconds. Display missing data rather than substituting a green zero.
3. Exercise failure and recovery
First observe a passing result and recent timestamp. Stop only the Ledger port-forward; this deliberately breaks the observer's access path, not the Ledger service itself. Observe failed checks and the alert after its configured delay. Restore the port-forward and inspect recovery. Keep this distinction in your incident record.
To learn notification routing, start a local Alertmanager with a no-delivery training receiver, add it to Prometheus' alerting.alertmanagers configuration, and reload/restart after validation. Confirm receipt in its local UI; a receiver with no integration does not send notifications. Actual external paging is outside this lab.
Checkpoint and optional telemetry
Save the observed metrics, dashboard JSON, failure/recovery times, and what the experiment actually affected. Stop the exporter and local telemetry processes when finished. Add Alloy/Loki and OpenTelemetry using their preceding bites as separate focused exercises; a combined stack is not required on the 16 GB host. Record those integrations as complete only after verifying ingestion and correlation.
Check the meaning of each observation
Before the first completion, the exporter's zero result is an initialization value. After a completed failure, it describes a real failed read. The timestamp distinguishes them. When you stop the Ledger port-forward, the backend can remain healthy while the observer loses access; write exactly that in the record. If up becomes zero instead, inspect whether the exporter itself stopped or its scrape path broke.
The simple rules can overlap during initialization or observer failure. That is an opportunity to refine tested routing/inhibition, not evidence that several independent outages occurred. The separate rule-test fixture in the alerting bite includes an action annotation; if testing this lab's shorter rule, match the expected annotations to this exact file. Finish with one explained signal path before adding optional collectors.
Sources
Prometheus installation↗, configuration↗, Grafana installation↗, Grafana server configuration↗, Alertmanager configuration↗, and Python HTTP server↗.
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.