Case study · Reliability Sprint
Cutting on-call load without adding headcount.
Deleted alert noise, tied every page to an SLO, and gave senior engineers their nights back.
The challenge
On-call had become a punishment shift, not a responsibility.
Twenty-seven pages a week sounds survivable until you notice most of them fired on symptoms nobody could act on at 3am — a disk nudging a threshold, a retry storm that resolved itself in ninety seconds. Engineers stopped trusting the pager, which meant the pages that mattered got the same shrug as the noise.
What we did
Rebuilt the alert set around what a human can actually do about it.
- Audited every alert against a real SLO. If a page didn’t map to something a customer would notice, it got deleted or downgraded to a dashboard.
- Rewrote the runbooks that survived. Every remaining alert now links to a runbook with the actual fix, not a wiki page from two reorgs ago.
- Gave on-call authority to silence noise on sight. No approval process to mute a bad alert — if it pages and doesn’t help, it’s gone by morning.
- Rebalanced the rotation. Fewer, more meaningful pages meant the rotation could shrink without anyone carrying more risk.
The result
Pages per on-call week dropped from 27 to 6 — and the six that remain get taken seriously, because they’re real. Median acknowledgment time fell from nine minutes to two, simply because engineers stopped assuming it was noise.
No new hires. Just a pager worth believing again.