Prometheus alerting rules that don't cry wolf
Write alerts on symptoms with for-durations and good labels so on-call gets pages that actually mean something.
An alert that pages on node_cpu above 80% for thirty seconds is technically correct and operationally useless: CPU is a cause, users do not feel it, and the on-call engineer who is woken by it learns to sleep through the next one. The alerts that hold up over years measure symptoms (error ratio, latency against the SLO, queue age) and carry enough context in their labels and annotations that the person paged can act without opening three dashboards first. Prometheus gives you the pieces for that; the rules file is where the discipline lives.
Four mechanisms do most of the work: a recording rule that names the ratio once, a for duration that keeps a blip from paging, a keep_firing_for duration that keeps a flapping condition from resolving and re-firing every evaluation, and labels that let Alertmanager decide who gets a page and who gets a ticket.
Record the ratio, then alert on the record
A rule like sum(rate(http_requests_total{status=~"5.."}[5m])) by (job) / sum(rate(http_requests_total[5m])) by (job) evaluated inside five alert rules is five copies of the same expression, five chances for them to drift, and five times the evaluation cost. A recording rule computes it once under a stable name that dashboards can share, and the naming convention level:metric:operations (job:http_requests:error_rate5m) tells the next reader the aggregation level, the source metric and what was done to it.
groups:- name: api_recordinginterval: 30srules:- record: job:http_requests:error_rate5mexpr: |sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)/sum(rate(http_requests_total[5m])) by (job)
for, keep_firing_for, and what pending means
for: 5m means the expression has to stay true across every evaluation for five minutes before the alert fires; between the first true evaluation and that point the alert is pending, visible in the UI and in ALERTS{alertstate="pending"} but not sent. Pick the duration from the metric's own noise: at least twice the scrape interval, and long enough that a deploy's brief error spike passes without a page. keep_firing_for: 5m addresses the opposite problem. Without it an alert resolves on the first evaluation where the condition is false, so an error ratio hovering around the threshold produces a resolved/firing pair every minute; with it the alert stays firing until the condition has been false for the full duration.
Labels are routing keys: severity decides page or ticket, team decides whose pager. Annotations are for humans, and they can use the alert's own labels and value through templating, so the notification says which job and how bad rather than "something is wrong". The runbook URL is not decoration; an alert whose runbook does not exist is a signal that nobody has decided what the response is.
groups:- name: api_alertsrules:- alert: HighErrorRateexpr: job:http_requests:error_rate5m > 0.05for: 5mkeep_firing_for: 5mlabels:severity: pageteam: platformannotations:summary: "5xx ratio {{ $value | humanizePercentage }} on {{ $labels.job }}"description: "More than 5% of requests to {{ $labels.job }} have failed for 5 minutes."runbook: "https://wiki.example.com/runbooks/high-error-rate"
Test the rule before it can page anyone
promtool check rules catches YAML and PromQL errors. promtool test rules catches the mistake that matters: a threshold or for that does not fire when it should, or fires when it should not. A unit test declares a synthetic series, an evaluation time, and the alerts expected at that time, so the for: 5m above can be proven pending at four minutes and firing at seven without touching a live Prometheus. Seven is a comfortably-firing checkpoint, not the moment it starts: with this input the alert is still pending at five minutes and firing by six, because for runs from the first evaluation the expression was true, one interval after the series begins.
rule_files:- ../recording-rules.yml- ../alert-rules.ymlevaluation_interval: 30stests:- interval: 30sinput_series:- series: 'http_requests_total{job="api", status="500"}'values: '0+10x40' # 10 errors per 30s, steadily- series: 'http_requests_total{job="api", status="200"}'values: '0+90x40' # 90 successes per 30s: a 10% error ratioalert_rule_test:- eval_time: 4malertname: HighErrorRateexp_alerts: [] # still pending: for is 5m- eval_time: 7malertname: HighErrorRateexp_alerts:- exp_labels: { severity: page, team: platform, job: api }exp_annotations:summary: "5xx ratio 10% on api"description: "More than 5% of requests to api have failed for 5 minutes."runbook: "https://wiki.example.com/runbooks/high-error-rate"
promtool check rules recording-rules.yml alert-rules.ymlChecking recording-rules.yml SUCCESS: 1 rules foundChecking alert-rules.yml SUCCESS: 1 rules foundpromtool test rules tests/api_alerts_test.yml # pending at 4m, firing at 7m SUCCESSpromtool test rules tests/api_alerts_test.yml # the same test asserting firing at 4m30s FAILED: alertname: HighErrorRate, time: 4m30s, exp:[ … HighErrorRate{severity="page", team="platform", job="api"} … ], got:[]the alert is still pending at 4m30s: pending through 5m, firing by 6m, so 7m is a checkpoint, not the startpromtool check rules broken.yml # expr with an unbalanced sum(rate(...) by (job) FAILED:broken.yml: group "broken", rule 1, "Bad": could not parse expression: 1:35: parse error: unexpected <by> in aggregationpromtool check config prometheus.yml SUCCESS: 2 rule files found SUCCESS: prometheus.yml is valid prometheus config file syntaxcurl -s localhost:9090/api/v1/rules | jq -r '.data.groups[].rules[] | "\(.name) \(.state // "recording")"'job:http_requests:error_rate5m recordingHighErrorRate inactivethe recording rule is always "recording"; the alert is "inactive" until its expression is true, then "pending", then "firing"Page or ticket is a label, not a feeling
Alertmanager routes on labels, so the page-versus-ticket decision has to be made in the rule. A page is for a condition a human must act on now: error-budget burn, a service unreachable, a data-loss risk, a security symptom such as a flood of 401s. Everything else, from a disk at 70% to a certificate expiring in two weeks, is a ticket that can wait for working hours. Multi-window burn-rate alerts, which page only when a short and a long window both show budget burn, are the refinement to add once the single-window rules have earned the on-call rotation's trust.
Grouping and inhibition are the Alertmanager side of the same idea: group_by: [alertname, job] collapses one incident into one notification, and an inhibition rule keeps HighErrorRate from paging alongside the PodRestarting alert that explains it. The Detection engineering course takes the rules from here into Alertmanager routing, Loki log correlation and the review loop that keeps a rules file honest.