Skip to content
← Advanced DevOps

Learning bite

Alertmanager and actionable alerts

Separate detecting a condition from routing a useful response.

Documentation reviewed2026-10-01 · 4 min read
On this page

Design an alert someone can act on

Prometheus evaluates an alert expression. A for interval requires a condition to remain true before it fires. Alertmanager groups, routes, silences, and inhibits notifications. These are separate responsibilities; a firing rule does not prove a notification reached anyone.

For this local path, keep notification delivery local: inspect Prometheus and Alertmanager, or route to a disposable local receiver. Do not send messages to real on-call channels during practice.

A useful MicroBank alert states the observed user-facing failure, the environment, how to inspect it, and a recovery runbook. Distinguish journey failure, exporter failure, and a stale check. A missing series is not a healthy value of zero.

Follow a condition through its states

An alert expression first needs a matching series whose condition is true. With for: 2m, that series is pending until the condition has stayed true across evaluations for two minutes; it can then become firing. If it stops matching first, the pending period resets. Alertmanager receives firing alerts and decides how to notify a receiver. Its grouping can combine related alerts; a silence mutes selected notifications for a period; inhibition can suppress dependent notifications while a higher-level alert is active. These actions do not change the underlying Prometheus measurement.

Work through two invented timelines at a 15-second evaluation interval. Failure first appears at 00:00 and disappears at 01:45: no firing alert should result. Failure remains true through 02:00: the two-minute condition can fire on that evaluation. Notification arrival can be later because of routing and grouping delays. A dashboard screenshot at 01:00 cannot establish that the notification path worked.

An actionable response distinguishes “the saved-case check failed” from “the exporter is unreachable” and “the last completion is too old.” The first points toward the checked path, the second toward the observer, and the third toward scheduling or a stalled check. They may share a cause, but that remains a hypothesis.

Write and test a small rule

This lab rule detects a last result that remains in a failed state. Its duration is a training choice, not a production SLO:

yaml
groups:
  - name: microbank-study
    rules:
      - alert: MicroBankJourneyFailing
        expr: microbank_journey_success == 0
        for: 2m
        labels:
          severity: warning
          environment: local
        annotations:
          summary: "Local MicroBank journey is failing"
          action: "Check probe logs, Accounts, queue, and Ledger; follow the local recovery record."

Save as infra/study/observability/rules.yaml; the module lab adds freshness and scrape rules. Use promtool check rules and a unit-test fixture before enabling it. Validate the negative case as well: a brief failure should not satisfy a two-minute for interval.

Test the rule without stopping a service

Install/select promtool with the same Prometheus release you plan to use. With rules.yaml saved as above, create infra/study/observability/rules-test.yaml:

yaml
rule_files:
  - rules.yaml
evaluation_interval: 15s
tests:
  - interval: 15s
    input_series:
      - series: microbank_journey_success
        values: '0+0x8 1'
    alert_rule_test:
      - eval_time: 1m
        alertname: MicroBankJourneyFailing
        exp_alerts: []
      - eval_time: 2m
        alertname: MicroBankJourneyFailing
        exp_alerts:
          - exp_labels:
              severity: warning
              environment: local
            exp_annotations:
              summary: "Local MicroBank journey is failing"
              action: "Check probe logs, Accounts, queue, and Ledger; follow the local recovery record."
      - eval_time: 2m15s
        alertname: MicroBankJourneyFailing
        exp_alerts: []

From that directory run promtool test rules rules-test.yaml. This fixture has failed samples from time zero through two minutes, then success at 2m15s. Expected: the test passes with no firing alert at one minute, a firing alert at two, and none after recovery. Change the samples to '0+0x6 1+0x2' and remove the expected two-minute alert; that failure ends at 1m45s and should never fire. Restore the first fixture if you want to keep both cases in separate test entries. This checks rule evaluation, not Alertmanager delivery.

Checkpoint

During a short, controlled test, record the timestamps for pending → firing → resolved. Explain why silencing a notification does not repair MicroBank, and why a health alert without a runbook becomes recurring noise.

Checkpoint answer: silence changes notification behavior, not the gauge or the application. If no series exists, metric == 0 alone cannot detect it; that needs a separately tested absence/scrape rule. Next, add contextual logs so an alert has somewhere useful to lead.

Sources

Prometheus alerting rules↗, rule tests↗, and Alertmanager↗.

Your notes and evidence

Record observations, questions, or links to your work. Keep credentials out of your notes.

Loading saved progress…

Back up or restore this path

Progress and notes stay in this browser. A backup contains only this learning path.