Prometheus alerting rules that don't cry wolf

Write alerts on symptoms with for-durations and good labels so on-call gets pages that actually mean something.

Sep 30, 2025·Updated ·7 min readIntermediate·By SecOpsLog · command-tested

An alert that pages on node_cpu above 80% for thirty seconds is technically correct and operationally useless: CPU is a cause, users do not feel it, and the on-call engineer who is woken by it learns to sleep through the next one. The alerts that hold up over years measure symptoms (error ratio, latency against the SLO, queue age) and carry enough context in their labels and annotations that the person paged can act without opening three dashboards first. Prometheus gives you the pieces for that; the rules file is where the discipline lives.

Four mechanisms do most of the work: a recording rule that names the ratio once, a for duration that keeps a blip from paging, a keep_firing_for duration that keeps a flapping condition from resolving and re-firing every evaluation, and labels that let Alertmanager decide who gets a page and who gets a ticket.

Record the ratio, then alert on the record

A rule like sum(rate(http_requests_total{status=~"5.."}[5m])) by (job) / sum(rate(http_requests_total[5m])) by (job) evaluated inside five alert rules is five copies of the same expression, five chances for them to drift, and five times the evaluation cost. A recording rule computes it once under a stable name that dashboards can share, and the naming convention level:metric:operations (job:http_requests:error_rate5m) tells the next reader the aggregation level, the source metric and what was done to it.

recording-rules.yml
groups:
- name: api_recording
interval: 30s
rules:
- record: job:http_requests:error_rate5m
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
/
sum(rate(http_requests_total[5m])) by (job)

for, keep_firing_for, and what pending means

for: 5m means the expression has to stay true across every evaluation for five minutes before the alert fires; between the first true evaluation and that point the alert is pending, visible in the UI and in ALERTS{alertstate="pending"} but not sent. Pick the duration from the metric's own noise: at least twice the scrape interval, and long enough that a deploy's brief error spike passes without a page. keep_firing_for: 5m addresses the opposite problem. Without it an alert resolves on the first evaluation where the condition is false, so an error ratio hovering around the threshold produces a resolved/firing pair every minute; with it the alert stays firing until the condition has been false for the full duration.

Labels are routing keys: severity decides page or ticket, team decides whose pager. Annotations are for humans, and they can use the alert's own labels and value through templating, so the notification says which job and how bad rather than "something is wrong". The runbook URL is not decoration; an alert whose runbook does not exist is a signal that nobody has decided what the response is.

alert-rules.yml
groups:
- name: api_alerts
rules:
- alert: HighErrorRate
expr: job:http_requests:error_rate5m > 0.05
for: 5m
keep_firing_for: 5m
labels:
severity: page
team: platform
annotations:
summary: "5xx ratio {{ $value | humanizePercentage }} on {{ $labels.job }}"
description: "More than 5% of requests to {{ $labels.job }} have failed for 5 minutes."
runbook: "https://wiki.example.com/runbooks/high-error-rate"

Test the rule before it can page anyone

promtool check rules catches YAML and PromQL errors. promtool test rules catches the mistake that matters: a threshold or for that does not fire when it should, or fires when it should not. A unit test declares a synthetic series, an evaluation time, and the alerts expected at that time, so the for: 5m above can be proven pending at four minutes and firing at seven without touching a live Prometheus. Seven is a comfortably-firing checkpoint, not the moment it starts: with this input the alert is still pending at five minutes and firing by six, because for runs from the first evaluation the expression was true, one interval after the series begins.

tests/api_alerts_test.yml
rule_files:
- ../recording-rules.yml
- ../alert-rules.yml
evaluation_interval: 30s
tests:
- interval: 30s
input_series:
- series: 'http_requests_total{job="api", status="500"}'
values: '0+10x40' # 10 errors per 30s, steadily
- series: 'http_requests_total{job="api", status="200"}'
values: '0+90x40' # 90 successes per 30s: a 10% error ratio
alert_rule_test:
- eval_time: 4m
alertname: HighErrorRate
exp_alerts: [] # still pending: for is 5m
- eval_time: 7m
alertname: HighErrorRate
exp_alerts:
- exp_labels: { severity: page, team: platform, job: api }
exp_annotations:
summary: "5xx ratio 10% on api"
description: "More than 5% of requests to api have failed for 5 minutes."
runbook: "https://wiki.example.com/runbooks/high-error-rate"
bash — observed: validate, test, and the state a wrong test catches (promtool 3.14.0)observed
promtool check rules recording-rules.yml alert-rules.yml
Checking recording-rules.yml
SUCCESS: 1 rules found
Checking alert-rules.yml
SUCCESS: 1 rules found
promtool test rules tests/api_alerts_test.yml # pending at 4m, firing at 7m
SUCCESS
promtool test rules tests/api_alerts_test.yml # the same test asserting firing at 4m30s
FAILED:
alertname: HighErrorRate, time: 4m30s,
exp:[ … HighErrorRate{severity="page", team="platform", job="api"} … ], got:[]
the alert is still pending at 4m30s: pending through 5m, firing by 6m, so 7m is a checkpoint, not the start
promtool check rules broken.yml # expr with an unbalanced sum(rate(...) by (job)
FAILED:
broken.yml: group "broken", rule 1, "Bad": could not parse expression: 1:35: parse error: unexpected <by> in aggregation
bash — observed: the rule states in a running Prometheus (no scrape target, so the alert has no data)observed
promtool check config prometheus.yml
SUCCESS: 2 rule files found
SUCCESS: prometheus.yml is valid prometheus config file syntax
curl -s localhost:9090/api/v1/rules | jq -r '.data.groups[].rules[] | "\(.name) \(.state // "recording")"'
job:http_requests:error_rate5m recording
HighErrorRate inactive
the recording rule is always "recording"; the alert is "inactive" until its expression is true, then "pending", then "firing"

Page or ticket is a label, not a feeling

Alertmanager routes on labels, so the page-versus-ticket decision has to be made in the rule. A page is for a condition a human must act on now: error-budget burn, a service unreachable, a data-loss risk, a security symptom such as a flood of 401s. Everything else, from a disk at 70% to a certificate expiring in two weeks, is a ticket that can wait for working hours. Multi-window burn-rate alerts, which page only when a short and a long window both show budget burn, are the refinement to add once the single-window rules have earned the on-call rotation's trust.

Alerts nobody acts on train people to ignore the ones that matter
An alert that fired in the last six months without a human doing anything is a candidate for deletion or for demotion to a ticket, and a cause-only alert (a pod restarted once) should page only if a symptom alert is also firing. Review ALERTS{alertstate="firing"} over the last month in the on-call retro; the list is usually shorter than the rules file suggests.
What was run for this article
promtool and Prometheus 3.14.0 (the prom/prometheus:v3.14.0 image, Docker Engine 28.5.2, linux/arm64) against the recording rule, alert rule and unit test shown above (build/evidence/prometheus-alerts). The observed blocks are copied from that run: check passes, the shipped unit test passes, the same test with a firing assertion at 4m30s fails, a rule with a broken expression fails check, and a running Prometheus reports the recording rule as recording and the alert as inactive. Seven assertions are checked; because the fixture has no metrics database and no scrape target, the results do not drift with time. Alertmanager routing, grouping and inhibition are the documented layer this stops short of.

Grouping and inhibition are the Alertmanager side of the same idea: group_by: [alertname, job] collapses one incident into one notification, and an inhibition rule keeps HighErrorRate from paging alongside the PodRestarting alert that explains it. The Detection engineering course takes the rules from here into Alertmanager routing, Loki log correlation and the review loop that keeps a rules file honest.

Related posts

Quick reference