Learning bite
Alertmanager and actionable alerts
Separate detecting a condition from routing a useful response.
On this page
Design an alert someone can act on
Prometheus evaluates an alert expression. A for interval requires a condition to remain true before it fires. Alertmanager groups, routes, silences, and inhibits notifications. These are separate responsibilities; a firing rule does not prove a notification reached anyone.
For this local path, keep notification delivery local: inspect Prometheus and Alertmanager, or route to a disposable local receiver. Do not send messages to real on-call channels during practice.
A useful MicroBank alert states the observed user-facing failure, the environment, how to inspect it, and a recovery runbook. Distinguish journey failure, exporter failure, and a stale check. A missing series is not a healthy value of zero.
Follow a condition through its states
An alert expression first needs a matching series whose condition is true. With for: 2m, that series is pending until the condition has stayed true across evaluations for two minutes; it can then become firing. If it stops matching first, the pending period resets. Alertmanager receives firing alerts and decides how to notify a receiver. Its grouping can combine related alerts; a silence mutes selected notifications for a period; inhibition can suppress dependent notifications while a higher-level alert is active. These actions do not change the underlying Prometheus measurement.
Work through two invented timelines at a 15-second evaluation interval. Failure first appears at 00:00 and disappears at 01:45: no firing alert should result. Failure remains true through 02:00: the two-minute condition can fire on that evaluation. Notification arrival can be later because of routing and grouping delays. A dashboard screenshot at 01:00 cannot establish that the notification path worked.
An actionable response distinguishes “the saved-case check failed” from “the exporter is unreachable” and “the last completion is too old.” The first points toward the checked path, the second toward the observer, and the third toward scheduling or a stalled check. They may share a cause, but that remains a hypothesis.
Write and test a small rule
This lab rule detects a last result that remains in a failed state. Its duration is a training choice, not a production SLO:
groups:
- name: microbank-study
rules:
- alert: MicroBankJourneyFailing
expr: microbank_journey_success == 0
for: 2m
labels:
severity: warning
environment: local
annotations:
summary: "Local MicroBank journey is failing"
action: "Check probe logs, Accounts, queue, and Ledger; follow the local recovery record."
Save as infra/study/observability/rules.yaml; the module lab adds freshness and scrape rules. Use promtool check rules and a unit-test fixture before enabling it. Validate the negative case as well: a brief failure should not satisfy a two-minute for interval.
Test the rule without stopping a service
Install/select promtool with the same Prometheus release you plan to use. With rules.yaml saved as above, create infra/study/observability/rules-test.yaml:
rule_files:
- rules.yaml
evaluation_interval: 15s
tests:
- interval: 15s
input_series:
- series: microbank_journey_success
values: '0+0x8 1'
alert_rule_test:
- eval_time: 1m
alertname: MicroBankJourneyFailing
exp_alerts: []
- eval_time: 2m
alertname: MicroBankJourneyFailing
exp_alerts:
- exp_labels:
severity: warning
environment: local
exp_annotations:
summary: "Local MicroBank journey is failing"
action: "Check probe logs, Accounts, queue, and Ledger; follow the local recovery record."
- eval_time: 2m15s
alertname: MicroBankJourneyFailing
exp_alerts: []
From that directory run promtool test rules rules-test.yaml. This fixture has failed samples from time zero through two minutes, then success at 2m15s. Expected: the test passes with no firing alert at one minute, a firing alert at two, and none after recovery. Change the samples to '0+0x6 1+0x2' and remove the expected two-minute alert; that failure ends at 1m45s and should never fire. Restore the first fixture if you want to keep both cases in separate test entries. This checks rule evaluation, not Alertmanager delivery.
Checkpoint
During a short, controlled test, record the timestamps for pending → firing → resolved. Explain why silencing a notification does not repair MicroBank, and why a health alert without a runbook becomes recurring noise.
Checkpoint answer: silence changes notification behavior, not the gauge or the application. If no series exists, metric == 0 alone cannot detect it; that needs a separately tested absence/scrape rule. Next, add contextual logs so an alert has somewhere useful to lead.
Sources
Your notes and evidence
Record observations, questions, or links to your work. Keep credentials out of your notes.
Back up or restore this path
Progress and notes stay in this browser. A backup contains only this learning path.