Observability & SRE interview questions
Practice observability and SRE interview answers from golden signals and PromQL through SLOs, tracing, alert hygiene, and incident response.
Levels run Beginner → Intermediate → Advanced → Expert. Answers are phrased the way you would say them in an interview; Advanced and Expert answers add the deeper reasoning, a diagram where it helps, and the follow-up an interviewer often asks next.
Fundamentals
Metrics are numeric time series, so they're what you alert on to detect that something's wrong. Traces follow a request across services and locate where it slowed down or failed. Logs are discrete events that explain what happened at that point. Together they tell you that something's wrong, where, and often why.
metrics → detect (error rate spike) traces → locate (slow dependency span) logs → explain (timeout message + fields)Link to this question
The short version is monitoring watches known failure modes with predefined alerts and dashboards. Observability is the property that lets you ask new questions about unknown failures from the telemetry you already have — without shipping new code.
monitoring: is CPU > 90%? observability: why are only Android users in region X slow?Link to this question
Latency, traffic, errors, and saturation — Google SRE's minimal set for user-facing systems. If I could only watch four things, these catch most of the user-impacting problems.
latency → p50/p99 request duration traffic → requests per second errors → 5xx / success ratio saturation→ queue depth, CPU, thread poolLink to this question
RED — Rate, Errors, Duration — is the user's view of a request-driven service. USE — Utilization, Saturation, Errors — is for resources like CPU, disk, and queues. I start with RED for user impact, then USE to find the bottleneck.
On an incident I usually ask 'are users hurting?' with RED before 'which resource is maxed?' with USE. Mixing them on one graph confuses responders. Microservices get RED dashboards per service; node and DB tiers get USE. Saturation — queue depth, thread pool wait — often predicts latency before utilization hits 100%.
api service: rate / error ratio / latency histogram postgres host: CPU util / runnable queue / disk errorsLink to this question
Cardinality is how many unique label combinations a metric has. Each combo is a time series — so unbounded labels like user_id or request_id explode memory, query cost, and scrape time.
# BAD — unbounded
http_requests_total{user_id="..."}
# GOOD — bounded
http_requests_total{route="/api/pay", code="500"}
# put user_id on logs/traces insteadLink to this questionI'd rather emit logs as key/value JSON so I can filter and aggregate instead of grepping free text. Consistent fields — service, level, trace_id — are what let you correlate with metrics and traces.
{"ts":"2026-07-24T10:20:01Z","level":"error",
"service":"api","trace_id":"a1b2c3",
"msg":"db timeout","dur_ms":812}Link to this questionMetrics & PromQL
Prometheus scrapes HTTP /metrics endpoints on a schedule — pull model — discovering targets via service discovery like Kubernetes or Consul. Push is mainly for short-lived jobs through the Pushgateway; Prometheus 3.x can also accept OTLP pushes when started with --web.enable-otlp-receiver.
# pull means Prometheus notices a dead target (failed scrape) # apps just expose /metrics; SD keeps targets current as Pods churnLink to this question
Counter is monotonic — requests_total. Gauge goes up and down — memory. Histogram buckets observations for quantiles; native histograms (exponential buckets in one series) are stable since Prometheus 3.8. Summary does client-side quantiles. The type decides which queries are even valid.
http_requests_total # counter node_memory_MemAvailable_bytes # gauge http_request_duration_seconds_bucket # histogramLink to this question
I'd divide the rate of 5xx-labeled requests by the rate of all requests over that window. rate() needs counters; sum() aggregates across instances.
sum(rate(http_requests_total{code=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))Link to this questionrate() is the per-second average over the whole range, so it's what I alert and build SLOs on. irate() uses only the last two samples, which suits zoomed-in graphs of volatile counters but is too jumpy for alerts. increase() is rate() times the range, for 'how many in the last hour' panels. All three adjust for counter resets, and increase() can return non-integers because it extrapolates.
rate(http_requests_total[5m]) # alerts, SLOs, recording rules irate(http_requests_total[5m]) # spiky dashboard view only increase(http_requests_total[1h]) # "requests in the last hour" # aggregate after: sum(rate(...)), never rate(sum(...))
Interviewer often follows with: Why should the rate() range cover several scrape intervals?
Link to this questionI'd use histogram_quantile over the rate of _bucket series. You can't average pre-aggregated averages and get a correct tail latency.
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m]))) # native histogram: no _bucket suffix, no by (le) histogram_quantile(0.99, sum(rate(http_request_duration_seconds[5m])))Link to this question
Native histograms are stable since Prometheus 3.8, but scraping them is opt-in with scrape_native_histograms. I'd enable it alongside always_scrape_classic_histograms so both forms exist, rewrite queries to drop _bucket and by (le), move SLO ratios to histogram_fraction, and only stop scraping classic buckets once dashboards, alerts, and the long-term store all read the native series.
A native histogram is one series with sparse exponential buckets, so you stop guessing bucket boundaries and cut series count, but every consumer has to understand the new sample type: recording rules, remote-write receivers, and the long-term store. Instrument with native support in the client library, enable scrape_native_histograms per job, and keep always_scrape_classic_histograms on during the overlap. Rewrite histogram_quantile to take rate(metric[5m]) directly and compute the SLO 'fast enough' ratio with histogram_fraction(0, 0.3, ...). Bound resolution with native_histogram_bucket_limit so a wide latency spread can't blow up memory. Compare old and new p99 and SLO ratios side by side for a full SLO window before cutting over.
scrape_configs:
- job_name: api
scrape_native_histograms: true
always_scrape_classic_histograms: true # keep _bucket during migration
native_histogram_bucket_limit: 160
histogram_quantile(0.99, sum(rate(http_request_duration_seconds[5m])))
histogram_fraction(0, 0.3, sum(rate(http_request_duration_seconds[5m])))Interviewer often follows with: What happens to a recording rule built on _bucket series once classic buckets stop being scraped?
Link to this questionI'd use recording rules to precompute expensive queries into new series for fast dashboards and reuse. Alerting rules evaluate conditions over time and fire to Alertmanager for grouping, routing, and silences.
- alert: HighErrorRate
expr: job:error_rate:ratio5m > 0.02
for: 10m
labels: { severity: page }
annotations: { summary: "5xx rate over 2%" }Link to this questionI'd page on user symptoms and SLO burn, such as error ratio and latency, rather than raw CPU. Route causes to tickets or dashboards, require every page to be urgent, actionable, and linked to a runbook, and delete or retune noise using pages-per-shift metrics.
Alert fatigue is a reliability bug. Start with an inventory: top alerts by volume and the share that led to action; delete flappers and demote always-firing alerts to tickets. Cause alerts like CPU and disk are fine on dashboards and for secondary diagnosis; primary pages should be symptoms tied to SLOs. Multi-window burn-rate alerts page on fast burns and ticket on slow ones. Alertmanager grouping and inhibition collapse storms during known outages. Every paging alert needs an owning team and a runbook link. I track MTTA/MTTR and pages that needed no action as hygiene KPIs, and review weekly until pages per night are back to a human number. Symptoms over causes, severity routing, and continuous deletion of noise.
page: SLO burn rate critical (5m+1h windows) ticket: disk > 80% with 2d forecast to full dashboard only: CPU steal, goroutine count
Interviewer often follows with: How would you introduce burn-rate alerts without double-paging during the migration?
Link to this questionYou alert on how fast the SLO error budget is being consumed across short and long windows — paging when a fast burn would exhaust the budget soon, not on every small spike.
The SRE workbook's multi-window, multi-burn-rate setup pages on 14.4× burn over 1h (5m short window) or 6× over 6h (30m short window), and opens a ticket on 1× over 3d (6h short window). That catches both 'site is on fire now' and 'we're quietly chewing budget.' It lines alerts up with the same SLI used for the SLO, so dashboards, pages, and policy share one definition of reliability. Tune windows to your SLO period — 30d is typical — and pair with a sensible for: and Alertmanager routes.
# page if burning 2% budget in 1h (approx) AND still burning over 5m # exact multipliers come from your SLO math / SRE workbook
Interviewer often follows with: What breaks if your SLI ignores a regional dependency that users feel?
Link to this questionI'd watch utilization and saturation — queue depth, thread wait, disk IO credit — under real traffic, load-test to find cliffs, and set autoscaling on leading indicators, not only CPU after users already hurt.
# scale on: concurrent requests or queue depth # not only: CPU > 90% after latency already burned SLOLink to this question
I'd identify top metrics by series count, drop unbounded labels, shard or remote-write to a long-term store, shorten local retention, and fix exporters at the source. Cardinality first, hardware second.
I'd use TSDB status and cardinality explorers to find offenders — things like http_requests_total{user_id=...}. Relabel-drop as an emergency brake, then patch apps and push per-request detail to logs and traces. Consider recording rules that aggregate away high-fanout labels for SLOs; on Kubernetes, be careful with pod and instance labels on high-fanout metrics and aggregate to deployment level where you can. Prefer histograms with bounded le buckets over per-instance summaries when you need global quantiles. Federate or use Thanos/Mimir/Cortex for scale, but don't remote-write garbage. Set series limits per tenant if you're multi-tenant, and capacity-plan series count, not only disk. Methodical cardinality triage beats buying bigger VMs.
1) top-N metrics by series 2) scrape relabel drops 3) fix instrumentation PRs 4) remote write + retention tiering
Interviewer often follows with: How would you attribute cardinality to a single team in a shared cluster?
Link to this questionLogging, tracing & cost
It's following one request across services as a trace of spans with parent/child links, propagated via context headers like W3C traceparent. It shows where latency and errors sit in the call graph.
# waterfall shows 400ms in payments client, not in the gateway # logs with the same trace_id explain the timeoutLink to this question
It's a vendor-neutral API, SDK, and collector for metrics, logs, and traces. You instrument once and export through a Collector — swapping backends becomes config, not a rewrite.
app SDK → OTLP → Collector (batch, sample, redact) → Jaeger/Tempo/vendorLink to this question
Apps send OTLP to a local agent Collector (DaemonSet or sidecar) for enrichment and batching. Agents forward to a tier running the load-balancing exporter keyed on trace ID, so every span of a trace lands on the same gateway Collector that runs tail sampling. memory_limiter goes first in every pipeline so overload becomes backpressure instead of OOM kills.
Tail sampling is stateful: the processor holds spans in memory until it can decide, and the docs are explicit that all spans of a trace must reach the same Collector instance. That forces a two-tier layout as soon as you run more than one sampler. The first tier uses the load-balancing exporter with a traceID routing key and DNS or static resolution of the second tier; the second tier runs tail sampling with policies for errors, latency over the SLO threshold, and a low probabilistic floor. Scale the second tier on memory and watch refused and dropped spans on both tiers. Metrics and logs don't need trace affinity, so route them through their own pipeline instead of the sampling tier.
app SDK --OTLP--> agent (DaemonSet): memory_limiter, k8sattributes, batch agent --> lb tier: loadbalancing exporter, routing_key: traceID lb tier --> sampler tier: memory_limiter, tail_sampling (errors, slow, 5% floor) sampler tier --> trace backend
Interviewer often follows with: What breaks if someone adds a second tail-sampling replica behind a plain round-robin Service?
Link to this questionI'd thread a shared trace_id through structured logs and spans, and use exemplars on latency metrics when they're available. Spike on a graph → exemplar or trace → logs for that id.
every log line: service, level, trace_id, span_id every outbound call: inject traceparentLink to this question
I'd use head sampling for cheap baseline traces and tail sampling to keep errors and slow outliers after the trace completes. Drop boring fast successes — never drop all 5xx just to save money.
Head sampling decides at the root span — cheap, but it may discard the one broken request. Tail sampling buffers spans in a Collector and decides with full knowledge of status, latency, and attributes, keeping the high-value traces. Every span of a trace must reach the same Collector instance, so at scale a load-balancing exporter tier sits in front of the samplers. Cost controls: attribute allow-lists, span limits, and separate retention for raw vs aggregated. Guaranteed throughput for error traces is a reliability feature. Document sampling so on-call knows a missing trace may be intentional.
keep: errors, latency > SLO threshold, canary traffic sample: 1% of successful GET /health-like traffic drop: high-cardinality custom attributes on spans
Interviewer often follows with: How does sampling interact with exemplars on Prometheus histograms?
Link to this questionI'd drop or sample debug and noisy sources, enforce structured fields with allow-lists, hash or omit unbounded keys, route debug to short retention, and fix chatty apps at the source. Cardinality in labels and indexes is as dangerous as raw GB/day.
Cost drivers are bytes ingested, indexed fields, and retention. Practical moves: default level=info in prod, sample successful access logs, never index user-generated strings as high-cardinality labels, and use dynamic rate limits per service. Keep full fidelity for errors and security audits. Metrics shouldn't duplicate every log line — metrics for aggregates, logs for examples. Fix the emitter; don't just buy a bigger plan.
# Collector/processor: drop http.access if code=200 and sample 1% # forbid labels: user_email, request_body # retention: debug 3d, info 14d, audit 365d
Interviewer often follows with: When is sampling access logs unacceptable for compliance?
Link to this questionRED metrics with histograms, structured logs with trace_id, OTel traces around outbound calls, a golden dashboard, SLO plus burn alerts, and cardinality guardrails — shipped as a platform template so every service starts the same.
Platform engineering wins here: auto-instrument HTTP/gRPC, a standard label vocabulary like service/route/code, log/metric/trace correlation baked in, and a service scaffold with /metrics, health probes, and dashboard JSON. Define which alerts are mandatory for production readiness. Don't let each team invent label schemes. Cost: default sampling and log levels. Readiness review checks: SLI defined, runbook linked from the alert, and on-call can find traces in under two minutes.
[ ] RED dashboard [ ] p99 + error-ratio burn alerts [ ] OTel trace on egress [ ] runbook URL on alert [ ] no unbounded metric labels
Interviewer often follows with: What do you require before the service gets a public DNS name?
Link to this questionSynthetics give controlled probes — uptime, multi-region path checks — even when traffic is low. Real-user and SLI metrics capture actual experience including client diversity. I alert on both: synthetics for 'is it up,' RUM/SLI for 'is it good for users.'
Synthetics can miss issues that only hit certain tenants, devices, or payloads; RUM can be delayed or sparse at night. Best practice for me: critical-path synthetics from multiple regions with auth where needed, plus SLIs from production traffic. I don't let a green synthetic dashboard override a burning user SLO. Keep synthetics out of the SLI denominator — or label and exclude them — so you don't inflate availability.
synthetic: every 60s HTTPS GET /login from 3 regions → page on fail SLI: real checkout success_fast ratio → burn alerts
Interviewer often follows with: How would you authenticate synthetics without creating a backdoor?
Link to this questionSLOs & alerting
An SLI is what you measure, say success ratio. An SLO is the internal target, like 99.9%. The error budget is 1 minus the SLO: the unreliability you're allowed to spend on change before you have to stabilize. An SLA is the external contract, usually looser than the SLO, with customer credits when you miss it.
SLI: successful HTTP requests / total SLO: 99.9% success over 30d budget: 0.1% ≈ 43 minutes of downtime / 30d SLA: 99.5% with customer credits budget spent → feature freeze / reliability workLink to this question
I'd pick something close to user experience: successful requests within a latency threshold — availability times freshness — not CPU. Count only user-facing critical routes, exclude synthetic health checks, and document exclusions.
Bad SLIs like CPU or pod restarts disconnect reliability policy from users. Good patterns: ratio of requests with code!~"5.." and latency under 300ms on checkout paths, or successful job completion for async workers. Separate SLOs per customer-critical journey if you need to. The window — 30d rolling vs calendar — changes burn math. Publish the PromQL next to the SLO so alerts and reports can't drift. Review quarterly; product changes invalidate SLIs.
sli = successful_fast / total_valid successful_fast: code=~"2.." and latency_bucket <= 0.3s exclude: route="/healthz", synthetic=true
Interviewer often follows with: How would you handle multi-endpoint services where one rare RPC is business-critical?
Link to this questionI'd trigger the agreed policy: freeze or slow feature launches, prioritize reliability work that restores the SLI, and review whether the SLO or the implementation is wrong. The budget is a negotiation tool, not a surprise punishment.
Mature orgs write the policy before the crisis: what happens at 50% and 100% budget burn, who decides exceptions, and how customer commitments interact. Exhaustion should redirect engineering capacity to toil reduction, dependency hardening, and alert/SLO quality — not endless firefighting without systemic fixes. Sometimes the SLO is too tight for the architecture; renegotiate explicitly with data. Exceptions for must-ship legal changes need compensating risk acceptance. Organizational design matters as much as PromQL.
at 50% burn: reliability backlog gets priority slots at 100%: feature freeze except P0 security/legal exception: VP product + SRE approve with expiry
Interviewer often follows with: How would you stop teams from gaming the SLI by shrinking what they measure?
Link to this questionIt groups related alerts, deduplicates, routes by labels to page vs ticket channels, and supports silences and inhibition so a parent outage doesn't spam every child alert.
severity=page → PagerDuty severity=ticket → Jira/Slack inhibit: InstanceDown suppresses HighLatency on same instanceLink to this question
I'd use for: pending windows, keep_firing_for so a brief dip doesn't resolve and re-fire the alert, rate over sufficiently long ranges, and hysteresis via burn-rate or separate raise/clear thresholds — never alert on a single noisy scrape.
- alert: HighErrorRate expr: job:error_rate:ratio5m > 0.02 for: 10m keep_firing_for: 5mLink to this question
If buckets skip your SLO threshold — say SLO is 300ms but buckets sit at 100ms and 500ms — histogram_quantile and success ratios get inaccurate. I'd align buckets with SLO boundaries and expected tails.
histogram_quantile linearly interpolates within a bucket, so wide gaps around the SLO create optimistic or pessimistic error; exemplars only attach sample trace IDs and do not interpolate. Redesign instrumentation with explicit boundaries at the SLO and common percentiles. Releasing a bucket change creates a series break — document it. This is a frequent production footgun in interviews.
# SLO: 300ms — include le="0.3" buckets: [0.05, 0.1, 0.3, 0.5, 1, 2.5]
Interviewer often follows with: Do you need the same buckets on every service?
Link to this questionI'd slice SLIs by region and critical journey and set regional SLOs or weighted budgets so a small region can't burn unnoticed and a large region can't hide a small one's pain.
Global averages hide localized failure — classic SRE footgun. Design: per-region success ratio, per mobile/web clients, and multi-window burn alerts on each slice. Traffic weights belong in reporting, not in erasing user pain. Document which SLO gates release. Dimensionality of SLIs is an architecture choice.
# alert: burn rate on sli:http_success{region="ap-south-1"}
# dashboard: SLO triangles per region, not one global numberInterviewer often follows with: How many SLO dimensions is too many for on-call to reason about?
Link to this questionIncidents & SRE practice
I'd detect and declare early, name an incident commander, mitigate first — rollback, shed load, failover — before deep root cause, communicate on a cadence, then resolve and run a blameless postmortem with owned actions.
Roles: IC decides, ops mitigates, comms handles customers and stakeholders, scribe keeps the timeline. Mitigation beats investigation while budget burns. Use feature flags and known rollback paths. Severity defines response time and audience. Afterward: restore SLO tracking, revoke emergency privileges, and open the postmortem within days. Calm structure under pressure beats hero debugging.
1) declare SEV and IC in channel 2) graph golden signals + last deploy 3) rollback / disable flag if strongly correlated 4) status update even if cause unknown
Interviewer often follows with: When do you stop rolling back and instead forward-fix?
Link to this questionFacts over fault: timeline, user impact, contributing factors, what went well and poorly, and concrete action items with owners and due dates. The goal is system learning — not blaming the human who touched the keyboard.
impact (users, duration, SLO burn) timeline (UTC) contributing factors (not "root cause" alone) action items: owner + date + tracked ticketLink to this question
Toil is manual, repetitive, automatable operational work that scales with service size. I'd track it, cap on-call toil, and invest engineering time to knock out the top sources — pagers that need human babysitting are a smell.
toil: hand-restarting pods, manual certificate renewals anti-toil: HPA, cert-manager, runbooks that are actually scriptsLink to this question
A runbook is the actionable guide for a specific alert: symptoms, dashboards, mitigation steps, escalation. Every page-worthy alert should link to one — alerts without runbooks train people to ignore pages.
annotations: summary: "Checkout error budget burning" runbook_url: https://wiki/runbooks/checkout-burnLink to this question
I'd suspect a wrong SLI — averages hiding tails — missing regional labels, client-side failures not in server metrics, or sampling bias. Check p99, per-region RED, synthetics, traces for errors, and RUM or app-store crash data.
Classic traps: success ratio excludes timeouts classified as client disconnects; load balancer 5xx not counted on app metrics; one shard failing behind a healthy average; CDN cache masking origin death for some paths. Expand telemetry scope before declaring user error. Chase the user's path, not the convenient dashboard.
1) p99 vs avg 2) break down by region/zone/version 3) ingress/LB codes vs app codes 4) client RUM / mobile crashlytics 5) recent deploys + feature flags
Interviewer often follows with: How can canary analysis false-pass on averages but fail users in one geo?
Link to this questionMTTD is mean time to detect; MTTR is mean time to recover. Good alerts minimize MTTD for user-impacting issues without creating so much noise that MTTR suffers from fatigue and lost trust.
too few alerts → high MTTD too many pages → high MTTR (ignored/slow)Link to this question
I'd scale on saturation or queue metrics — or a latency SLI — that lead the incident, not only CPU, and expose those metrics from the app runtime.
CPU can be low while request queues grow — blocked threads, connection pools. USE/RED: watch pool utilization, queue depth, and p99. Custom metrics for HPA or KEDA. Load-test to find the right signal. Leading indicators tied to the failure mode beat CPU after users already hurt.
# metric: app_queue_depth > 50 → scale out # verify: CPU may stay at 20% while latency climbs
Interviewer often follows with: How would you prevent flapping when scaling on noisy queue metrics?
Link to this questionI'd emit deploy events as annotations or changelog markers, label metrics with version, retain traces and logs for the change window, and store a durable audit of who promoted which digest.
Change intelligence: deploy markers in Grafana, version labels on RED metrics, CI identity in provenance, and immutable artifact digests. Incident reviews should pull the marker timeline automatically. Retention must meet audit windows. Observability plus CI identity plus change calendar as one system.
http_requests_total{version="sha-abc"}
# Grafana annotation: deploy sha-abc by oidc:alice at TInterviewer often follows with: How would you attribute a canary's SLO burn to a specific digest under shared pods?
Link to this questionReal-world scenarios
I'd emergency-drop the high-cardinality labels via scrape relabel, restore TSDB health, then patch instrumentation and add cardinality gates in CI and review.
Unbounded labels create a series per combination and can OOM Prometheus. Mitigate with metric_relabel_configs drop/keep, restart if corrupted, and find the offender via cardinality explorers or tsdb status. Fix apps to put IDs on logs and traces. Prevent with lint rules and per-tenant series limits. Stabilize the platform before blaming 'just add RAM.'
metric_relabel_configs:
- regex: "user_id|request_id"
action: labeldrop
# labeldrop must leave series unique; if it doesn't, drop the metric instead
# then: PR removing those labels from the exporterInterviewer often follows with: Why can recording rules make a cardinality problem worse if written carelessly?
Link to this questionI'd check window length vs outage duration, the SLI definition — success vs availability — traffic weighting, and whether the burn threshold is too loose for short complete failures.
Classic miss: long windows dilute a short 100% outage; or the SLI counts synthetic checks that still pass while users fail; or the budget is so large a total outage barely burns. Fix: fast and slow burn pairs — such as the workbook's 1h/5m and 6h/30m page pairs plus a 3d/6h ticket pair — SLIs that match user journeys, and alert on absolute error spikes as a backup. Validate with historical incident replay. Distrust a green burn graph during known pain.
# did 30m at 0% success burn enough of the 30d budget to trip 2% / 1h? # add: page on success ratio < 99% for 5m OR fast burn # SLI: browser RUM or gateway 5xx — not only /healthz
Interviewer often follows with: When would a pure error-budget burn alert be the wrong primary page?
Link to this questionI'd move to tail-based sampling that keeps errors and high-latency traces at high rate, keep a low probabilistic sample for the happy path, and link metrics exemplars to retained traces.
Head sampling randomly drops the rare bad requests you need. Tail sampling in the collector inspects completed traces and retains on status/latency rules with a daily budget. Pair with structured logs carrying trace_id for the rest. Cost control: per-service caps and defensive drop of chatty health checks. Purposeful sampling beats uniform 1%.
# keep 100% if error OR p99-class latency # else 1% probabilistic # budget: max N spans/day/service
Interviewer often follows with: What bias appears if you keep 100% of errors during a dependency flap?
Link to this questionI'd stop the leak — scrub or redact at source or ingest — rotate exposed sessions and tokens, restrict who can query historical logs, and add CI/lint rules that block PII fields going forward.
Logs are a data store under compliance regimes. Immediate: deploy redaction processors, delete or quarantine hot indexes if policy requires it, rotate credentials, and notify privacy/security. Lasting: field allowlists, structured logging standards, and DLP-ish scanners in CI. Never 'we'll clean it next sprint' while tokens remain valid.
# collector: replace email/token fields with hash or drop # app: log user_id (internal) not email; never Authorization headers # CI: fail on patterns Authorization|password|ssn
Interviewer often follows with: Which is safer for correlation — hashing emails in logs or using opaque internal user IDs?
Link to this questionI'd look at user-journey SLIs — checkout success, RUM, synthetics — not host CPU panels. Then traces and logs for the failing step, and ask which dependency the dashboards don't cover.
Green host metrics with broken UX means you instrumented the wrong layer, or health checks that don't exercise checkout. Path: confirm from the edge — CDN/gateway status, synthetic checkout — identify the failing span, check canaries and feature flags, and only then open the USE resource graphs. Fix the dashboard set to include journey RED. Dashboards lie when SLIs are vanity.
1) synthetic checkout / RUM conversion 2) gateway 5xx + latency on /checkout/* 3) trace: payment span errors 4) only then: CPU/disk of payment pods
Interviewer often follows with: Give an example of a health check that stays 200 while checkout is down.
Link to this questionTonight I'd stabilize with first principles — RED, recent deploys, dependencies. Tomorrow I'd block paging alerts without runbooks and write the missing doc from the incident timeline.
Missing runbooks extend MTTR. Night-of: declare IC, capture timeline, use standard playbooks — rollback, shed load, failover. Next day: every alert rule requires annotations.runbook_url; PR checks reject pages without it; improve the doc with the commands that actually worked. Game days keep runbooks honest. An alert without a next step is unfinished work.
annotations: summary: "Checkout error budget fast burn" runbook_url: "https://runbooks/checkout-burn" dashboard_url: "https://grafana/.../checkout"
Interviewer often follows with: What belongs in a runbook vs a full architecture doc?
Link to this questionI'd recalibrate SLI quality and thresholds with incident replay, separate ticket vs page burns, show precision metrics, and only then restore paging — with a named owner for the SLO.
False SEVs usually mean bad SLIs — healthz, low-traffic noise — or windows that flap. Work: replay last month's tickets against proposed alerts; raise continuity requirements; exclude known maintenance; document when to page vs ticket. Publish alert precision weekly. Deleting SLO paging without replacement returns you to CPU roulette. This is socio-technical repair of the SLO program.
# for each past incident: would fast/slow burn have fired? T/F positive? # target: >80% precision before re-enabling page # owner: sre-checkout@
Interviewer often follows with: How would you handle a correct burn alert that the business still considers a false SEV?
Link to this questionI'd check remote-write queues, shard limits, and relabel drops on the write path — and treat missing samples as an observability SEV because under-reporting hides burns.
Silent remote-write failure is dangerous optimism. Symptoms: queue capacity alerts, 4xx/5xx from the backend, hashmod imbalance, or overly aggressive write relabel. Compare edge gateway success rates to stored SLI rates. Mitigate: buffer, spill to local, page on write failure, and never compute compliance-only SLOs solely from a lossy store. Telemetry pipelines need SLOs too.
prometheus_remote_storage_samples_failed_total prometheus_remote_storage_queue_highest_timestamp_seconds # compare: edge success ratio vs recording rule in long-term store
Interviewer often follows with: Why might a local Prometheus graph disagree with the global Thanos/Mimir view during a write outage?
Link to this questionI'd slice RED and SLIs by version or canary label and sticky pool, alert on per-version error rates, and verify cutover with traffic percent and unique version counts — not only aggregate green graphs.
Aggregates hide split-brain deploys. Instrument version labels on metrics, show dual burn charts during cutover, and run synthetic checks against both pools. Confirm Service/Gateway weights and session affinity behavior. Rollback criteria must include 'any version still serving with elevated errors,' not global averages.
http_requests_total{version="blue|green",code=~"5.."}
# alert: burn(version=green) OR (traffic_green>95% AND blue_errors high sticky remnant)Interviewer often follows with: What label cardinality trade-off do you accept to get per-version SLIs?
Link to this questionI'd normalize on trace_id and exemplars, check NTP and clock skew between nodes and SaaS, and treat one signal as primary for time while using others for detail — documenting skew in the timeline.
Skew and different batching windows create ghost causality. Practices: require trace_id in logs, exemplars from metrics to traces, understand scrape/timestamp configs, and note processing delay in log pipelines. For IR, prefer gateway or RUM timestamps as user-truth. Fix chronic skew — NTP — as a platform item. Correlation is a designed contract, not hope.
log: {"trace_id":"...","ts":"..."}
metric exemplar → trace
# IR note: log ingest delay ~45s; prefer gateway tsInterviewer often follows with: How would you detect systematic clock skew across a Kubernetes node pool?
Link to this questionI'd ship tiered runbooks, a single 'start here' dashboard with journey SLIs, auto-links from alerts, and a shadow/onboarding rotation — paging alerts without those artifacts don't ship.
Human reliability is part of the system. Design: alert → dashboard → runbook → rollback one-liner; dependency map; clear escalation. Keep runbooks short and command-oriented. Practice with game days. Reduce tribal knowledge by encoding it in alert annotations. MTTR is a product of docs and design, not heroics.
alert annotations → grafana journey board → runbook runbook: 1) confirm user impact 2) last deploy 3) rollback 4) escalate game day: quarterly checkout burn drill
Interviewer often follows with: What's the failure mode of a 40-page runbook during a SEV-1?
Link to this questionRelated
Primary references
Found a technical issue on this page? Report it with the tool version you used and the behavior you saw. How resources are maintained.