On-Call Runbook

What to check at 3 a.m. before you escalate. One runbook per critical service or high-volume alert — linked from monitors and on-call docs.

Symptoms → diagnostics → mitigations → escalate — with Incido fields baked in.

2 min read
On this page

Preview or copy each section below — paste into Notion, Confluence, or your wiki.

Service runbook

Runbook: [Service / alert]

Service card

FieldValue
Owner[Team]
On-call schedule[Schedule name]
Escalation policy[Policy name]
Status page component[Component name]
Monitors[Links]
Last testedYYYY-MM-DD

What it does

[1–2 sentences — why a page here matters]

If you see this

  • [Alert condition or symptom]
  • [User-visible behavior]

Severity hint

ConditionLevel
[All users down]Critical
[Elevated errors > X%]High
OtherwiseMedium — see severity matrix

Diagnose (in order)

  1. [Dashboard/monitor] — check [metric]
  2. Recent deploys — [CI/release link]
  3. Dependencies — [DB, queue, vendor status]
  4. Logs — [query]

Mitigate

StepActionSafe?Expected result
1[Restart workers / scale]YesErrors drop
2[Rollback / feature flag]Needs approvalService recovers
3Escalate after [X] minBring in [expert/secondary]

Rollback

[Link or exact steps · who approves]

If customer-visible

Publish to [component] on status page:

We are investigating [symptom] affecting [component].

After resolution

  • Timeline notes in incident
  • Follow-up for missing monitors/docs
  • Update this runbook

Known false positives

  • [Looks like outage but is not — how to tell]

How to use this template

  • One runbook per noisy alert or tier-1 service — not one mega-doc.
  • Test during game days; stale commands waste minutes.
  • Link from monitor runbook field or team wiki.

Open an incident from the alert, follow this runbook, log on the timeline, and publish status updates from the same record. Explore incident management →

FAQs

Runbook vs playbook?

Runbook = technical steps for one service. Playbook = roles and process for any incident.

Update cadence?

After any incident where steps were wrong or missing. Quarterly minimum for tier-1 services.

Include shell commands?

Yes for safe, repetitive ops. Mark destructive commands and note required approvals.

Put your process into practice

14-day free trial. On-call, incidents, and status pages — no credit card required.