Glossary

Incident management glossary

Clear definitions for the terms engineers use during incidents — with examples, common mistakes, and links to guides and templates.

Jul 12, 2026 · 2 min read

MTTR (Mean Time to Recovery)

MTTR is the average time from when an incident starts until service is restored to an acceptable level — a core reliability metric for measuring how fast your team recovers from outages.

Read definition
Jul 12, 2026 · 2 min read

MTTD (Mean Time to Detect)

MTTD is the average time from when an incident actually begins until your team recognizes it — measuring how quickly monitoring, alerts, and on-call practices surface real production problems.

Read definition
Jul 12, 2026 · 2 min read

SLO (Service Level Objective)

An SLO is a target level of reliability for a service — expressed as a percentage or threshold over a time window — that tells your team how much failure budget you can spend before customer trust is at risk.

Read definition
Jul 12, 2026 · 2 min read

SLI (Service Level Indicator)

An SLI is a quantifiable measure of how well your service is performing for users — such as availability, latency, or error rate — that forms the foundation for SLOs and incident severity decisions.

Read definition
Jul 12, 2026 · 2 min read

Incident Commander

The incident commander is the person responsible for coordinating response during an active incident — assigning roles, driving decisions, and keeping communication structured while engineers focus on mitigation.

Read definition
Jul 12, 2026 · 2 min read

Runbook

A runbook is a step-by-step guide for responding to a known operational scenario — giving on-call engineers concrete actions instead of improvising from memory during an incident.

Read definition
Jul 12, 2026 · 2 min read

Escalation Policy

An escalation policy defines who gets notified, in what order, and after which timeouts when an alert is not acknowledged — ensuring incidents reach a human who can act before impact grows.

Read definition
Jul 12, 2026 · 2 min read

Alert Fatigue

Alert fatigue is the desensitization that happens when engineers receive too many low-signal or false-positive pages — leading to missed acknowledgments, slower detection, and on-call burnout.

Read definition
Jul 12, 2026 · 2 min read

Status Page

A status page is a public (or subscriber-facing) site that communicates current service health, active incidents, and scheduled maintenance — reducing support load and building trust during outages.

Read definition
Jul 12, 2026 · 2 min read

On-Call Rotation

An on-call rotation is a repeating schedule that assigns engineers as first responders for production issues — balancing coverage, fairness, and clear handoffs so alerts always reach someone accountable.

Read definition

Ready to run incidents in one place?

Incido brings on-call, structured workflows, and status pages together.