Glossary
Incident management glossary
Clear definitions for the terms engineers use during incidents — with examples, common mistakes, and links to guides and templates.
MTTR (Mean Time to Recovery)
MTTR is the average time from when an incident starts until service is restored to an acceptable level — a core reliability metric for measuring how fast your team recovers from outages.
Read definitionMTTD (Mean Time to Detect)
MTTD is the average time from when an incident actually begins until your team recognizes it — measuring how quickly monitoring, alerts, and on-call practices surface real production problems.
Read definitionSLO (Service Level Objective)
An SLO is a target level of reliability for a service — expressed as a percentage or threshold over a time window — that tells your team how much failure budget you can spend before customer trust is at risk.
Read definitionSLI (Service Level Indicator)
An SLI is a quantifiable measure of how well your service is performing for users — such as availability, latency, or error rate — that forms the foundation for SLOs and incident severity decisions.
Read definitionIncident Commander
The incident commander is the person responsible for coordinating response during an active incident — assigning roles, driving decisions, and keeping communication structured while engineers focus on mitigation.
Read definitionRunbook
A runbook is a step-by-step guide for responding to a known operational scenario — giving on-call engineers concrete actions instead of improvising from memory during an incident.
Read definitionEscalation Policy
An escalation policy defines who gets notified, in what order, and after which timeouts when an alert is not acknowledged — ensuring incidents reach a human who can act before impact grows.
Read definitionAlert Fatigue
Alert fatigue is the desensitization that happens when engineers receive too many low-signal or false-positive pages — leading to missed acknowledgments, slower detection, and on-call burnout.
Read definitionStatus Page
A status page is a public (or subscriber-facing) site that communicates current service health, active incidents, and scheduled maintenance — reducing support load and building trust during outages.
Read definitionOn-Call Rotation
An on-call rotation is a repeating schedule that assigns engineers as first responders for production issues — balancing coverage, fairness, and clear handoffs so alerts always reach someone accountable.
Read definitionReady to run incidents in one place?
Incido brings on-call, structured workflows, and status pages together.