Alert Fatigue

Alert fatigue is the desensitization that happens when engineers receive too many low-signal or false-positive pages — leading to missed acknowledgments, slower detection, and on-call burnout.

2 min read
On this page

Alert fatigue is what happens when on-call engineers are overwhelmed by pages — too many, too noisy, or too often non-actionable — until they slow down, mute channels, or stop trusting alerts altogether.

Fatigue is a reliability problem, not a discipline problem. When every alert feels like crying wolf, real incidents hide in the noise and MTTD suffers.

Why it matters

Burned-out on-call engineers miss pages, delay triage, and leave the rotation. Teams then “fix” fatigue by adding more people to every notification, which makes fatigue worse. Breaking the cycle requires better signals, tighter escalation policies, and rotation design that shares load fairly — central themes in How to Build an Effective On-Call Rotation guide.

Example

A team pages on-call 40 times per week. Most alerts auto-resolve within two minutes or reflect known flaky tests. Responders start acknowledging without investigation. When a genuine database incident pages at 02:00, acknowledgment is delayed 12 minutes because the primary assumed it was noise. Effective detection lag — and customer impact — grows.

Common mistakes

  • Paging on warnings instead of user-impacting SLI breaches.
  • No alert ownership — orphaned monitors nobody will tune.
  • Duplicate alerts for the same root cause across ten dashboards.
  • Rewarding “always on” instead of fixing the sources of noise.

Treat alert volume as a metric: track pages per on-call shift, group by service, and delete or fix the noisiest monitors first. Every page should link to a runbook or explicit triage steps. If an alert has not required human action in 90 days, downgrade or remove it.

Healthy on-call is boring on-call — boring means your signals match real work.

Run incidents with structure

Incido brings on-call schedules, escalation policies, structured workflows, and status pages together.