# Engineering Incident Toolkit

Free copy-ready templates for incident response, timelines, customer communication, on-call, escalation, and runbooks. Adapt for your team — use in Notion, Confluence, Google Docs, or your wiki.

Source: https://incido.app/templates

---

## Table of contents

1. [Incident response playbook](#incident-response-template)
2. [Incident timeline template](#incident-timeline-template)
3. [Status page messages](#status-page-update-template)
4. [Maintenance planner](#maintenance-window-template)
5. [On-call schedule planner](#on-call-schedule-template)
6. [Escalation policy designer](#escalation-policy-template)
7. [Severity matrix](#severity-levels-template)
8. [On-call runbook](#runbook-template)

---

## 1. Incident response playbook {#incident-response-template}

A one-page playbook for when alerts fire. Maps directly to Incido’s incident stages — Triage, Active, Post Incident, and Closed — so responders know what to do and when to go public.

### How to use

- Pin this in your wiki and link from on-call runbooks.
- Align severity names and escalation policies with your Incido org settings.
- Run a tabletop quarterly — walk one scenario through each stage.

### Response playbook

# Incident response playbook

> Use from first alert through closure. Pair with the timeline template to record what happened.

## Incident card
| Field | Value |
|-------|-------|
| Title | [Short customer-facing name] |
| Deduplication key | [stable-key — same key merges in Triage/Active] |
| Stage | Triage → Active → Post Incident → Closed |
| Severity | Critical / High / Medium / Low |
| Affected components | [Component A, Component B] |
| Incident commander | [Name] |
| Comms lead | [Name] |
| Escalation policy | [Policy name] |

---

## Stage checklist

### Triage — assess before you broadcast
- [ ] Confirm new incident (check deduplication key / related alerts)
- [ ] Set severity and affected components
- [ ] Assign commander and comms lead
- [ ] Decide: status-page update needed now? **Yes / No**

### Active — confirmed impact, full response
- [ ] Move to **Active** stage
- [ ] Page on-call via escalation policy
- [ ] Log key decisions on the timeline
- [ ] Publish status update if customer-visible
- [ ] Notify subscribers on meaningful stage changes

### Post Incident — impact ended, follow-up remains
- [ ] Move to **Post Incident** when customers are no longer affected
- [ ] Publish resolved update on status page
- [ ] Capture follow-up tasks on the timeline

### Closed — wrapped up
- [ ] Follow-up tasks owned (not necessarily done)
- [ ] Move to **Closed**

---

## Roles
| Role | Owns |
|------|------|
| **Incident commander** | Coordination, decisions, timeline quality |
| **Comms lead** | Status page copy, subscriber updates, timing |
| **Technical lead** | Investigation, mitigation, recovery evidence |

## When to escalate manually
- Wrong on-call paged → **Bring in** a user, team, or schedule on the incident
- No acknowledgement → escalation policy advances automatically
- Scope grows → raise severity and notify leadership per severity matrix

## Customer update cadence
| Severity | First public update | Ongoing |
|----------|--------------------|---------|
| Critical | ≤ 15 min | Every 30–60 min while impacted |
| High | ≤ 30 min if visible | Every 60 min |
| Medium | If visible | On material changes |
| Low | Rarely | N/A |


---

## 2. Incident timeline template {#incident-timeline-template}

The timeline is your incident’s source of truth. Log facts as they happen — Incido also records pages, acknowledgements, and status updates automatically.

### How to use

- Assign one log owner during the incident — often the commander or a scribe.
- Record public updates when sent, not when drafted.
- Track follow-ups here; close the incident once tasks are owned.

### Timeline log

# Incident timeline

> Facts, not chat logs. One line per meaningful event.

## Header
| Field | Value |
|-------|-------|
| Incident | [Title] |
| Deduplication key | [key] |
| Timezone | [Europe/Zurich] |
| Log owner | [Name — keeps this current during the incident] |

---

## Live log
| Time | Who | Type | What happened |
|------|-----|------|---------------|
| HH:MM | Monitor | Detection | Alert: [name] — [symptom] |
| HH:MM | [Name] | Triage | Declared. Severity [X]. Components: [list] |
| HH:MM | System | Escalation | Paged [policy] → [Name] acknowledged |
| HH:MM | [Name] | Investigation | Hypothesis: [cause]. Checking [metric/deploy] |
| HH:MM | [Name] | Decision | Roll back [change] |
| HH:MM | [Name] | Public update | Status page: Investigating — "[summary]" |
| HH:MM | [Name] | Mitigation | [Action] → [result] |
| HH:MM | [Name] | Recovery | [Metric/test] normal |
| HH:MM | [Name] | Public update | Status page: Resolved — "[summary]" |
| HH:MM | [Name] | Wrap-up | → Post Incident. Follow-ups: [list] |

### Entry types
`Detection` · `Triage` · `Escalation` · `Investigation` · `Decision` · `Mitigation` · `Public update` · `Recovery` · `Wrap-up`

### Log these
Severity changes · Pages & acks · Published status updates · Mitigations & outcomes · Handoffs

### Skip these
Play-by-play chat · Unlabeled speculation · Secrets & customer PII

---

## Follow-up (after impact ends)
| Task | Owner | Due |
|------|-------|-----|
| [Add monitor / fix runbook] | [Name] | YYYY-MM-DD |


---

## 3. Status page messages {#status-page-update-template}

Calm, consistent wording reduces support load. Publish these from your incident or maintenance record — component health and subscriber email stay in sync.

### How to use

- Adapt tone to your brand — keep sentences short.
- Pair each message with the right component status (degraded, outage, operational).
- Use subscriber email for investigating → identified → resolved on Critical/High.

### Incident messages

Match component status on your status page to each message.

# Status page — incidents

## At a glance
| Field | Set in Incido |
|-------|---------------|
| Affected components | [list] |
| Severity | Critical / High / Medium / Low |
| Notify subscribers | On for Critical/High stage changes |

---

## Message templates

**Investigating**
> We are investigating [symptom] affecting [component(s)]. Next update by [time] or sooner.

**Identified**
> We have identified an issue affecting [component(s)] and are working on a fix. [One sentence on customer impact.]

**Monitoring**
> A fix has been deployed. We are monitoring [component(s)] and will confirm when fully resolved.

**Resolved**
> The incident affecting [component(s)] is resolved. All systems are operational.

**Resolved + follow-up**
> The incident affecting [component(s)] is resolved. We are reviewing what happened and will share customer-relevant findings if needed.

### Writing rules
1. Lead with **customer impact**, not internal cause
2. Set a **next update time** while investigating
3. Update **component status** to match reality
4. Notify subscribers on **material** changes only

### Maintenance messages

# Status page — maintenance

## At a glance
| Field | Value |
|-------|-------|
| Affected components | [list] |
| Window | [start] – [end] ([timezone]) |

---

## Message templates

**Scheduled**
> Maintenance on [component(s)] is scheduled from [start] to [end] ([timezone]). [Expected impact in plain language.]

**In progress**
> Maintenance on [component(s)] is in progress. [Expected impact]. We will update when complete.

**Extended**
> Maintenance on [component(s)] will take longer than planned. We now expect completion by [revised time].

**Complete**
> Maintenance on [component(s)] is complete. All systems are operational.

## Subscriber checklist
- [ ] Advance notice before window
- [ ] Start notification (if impact begins)
- [ ] Completion notification


---

## 4. Maintenance planner {#maintenance-window-template}

Plan maintenance before production day. Incido handles the window, affected components, optional incident suppression, and status-page publishing.

### How to use

- Create the record when the window is approved — not day-of.
- Map components to your status page before scheduling.
- Enable suppression only for expected monitor noise during the window.

### Maintenance plan

# Maintenance: [Title]

> Planned work with a clear customer story and rollback path.

## Plan
| Field | Value |
|-------|-------|
| Deduplication key | [stable-key] |
| Owner | [Name] |
| Window | [start] → [end] ([timezone]) |
| Auto-start / auto-end | Yes / No · grace [X] min |
| Suppress auto-incidents | Yes / No — only if monitors would false-page |
| Status page | [Page name] |

## Summaries
**Customers see:** [Plain language — e.g. brief disconnects, read-only mode]

**Team sees:** [Deploy steps, blast radius, rollback owner]

## Affected components
| Component | Status during window | Notes |
|-----------|---------------------|-------|
| [API] | Degraded | Rolling restart |
| [Dashboard] | Operational | No impact |

## Comms timeline
| When | Action |
|------|--------|
| [X]h before | Advance notice + subscriber email |
| Start | “In progress” update |
| End | “Complete” update |

### Copy-paste messages
**Advance:** Scheduled maintenance on [component] from [start] to [end]. [Impact sentence.]

**In progress:** Maintenance on [component] is in progress. [Impact]. We will update when done.

**Complete:** Maintenance on [component] is complete. All systems operational.

## Execution
**Before**
- [ ] Approval / change ticket recorded
- [ ] On-call aware
- [ ] Rollback documented
- [ ] Dashboards open

**During**
- [ ] Post in-progress at actual start
- [ ] Abort if: [metric/symptom] → rollback

**After**
- [ ] Verify key journeys
- [ ] Post complete + close maintenance
- [ ] Log issues on incident timeline if needed

## Rollback
[Steps · owner · time limit]


---

## 5. On-call schedule planner {#on-call-schedule-template}

Fair rotations and visible coverage prevent gaps. Document how your schedule connects to escalation policies before the first missed page.

### How to use

- Target ≤ 33% on-call load per engineer — adjust rotation length to team size.
- Keep overrides visible; surprise handoffs cause missed pages.
- Enable shift verification to catch bad contact methods early.

### Schedule spec

# On-call: [Team name]

## Overview
| Field | Value |
|-------|-------|
| Schedule | [Name in Incido] |
| Timezone | [Europe/Zurich] |
| Escalation policy | [Policy name] |
| Shift verification | On / Off |
| Last reviewed | YYYY-MM-DD |

## Rotation layers
| Layer | Length | Handoff | Rotation order |
|-------|--------|---------|----------------|
| Primary | 1 week | Mon 09:00 | A → B → C → D |
| Secondary | 1 week | Mon 09:00 | E → F → G |

## Coverage snapshot
| Week of | Primary | Secondary |
|---------|---------|-----------|
| YYYY-MM-DD | [Name] | [Name] |

## Handoff (5 min)
**Outgoing:** Open incidents · Upcoming maintenance · Contact methods verified

**Incoming:** Acknowledge shift · Test push/SMS/email · Know escalation policy name

## Overrides
| Dates | Covering | Reason |
|-------|----------|--------|
| YYYY-MM-DD – YYYY-MM-DD | [Name] | Vacation |

## Scope
**Covers:** [Services / components]

**Does not cover:** [Other teams’ components — link their schedule]

## Incident sources
| Source | How it reaches Incido |
|--------|----------------------|
| Monitors | API / webhook |
| Manual | Dashboard |
| Support | [Triage process] |


---

## 6. Escalation policy designer {#escalation-policy-template}

Who gets paged, in what order, and when to give up on step one. Design policies here, then configure them in Incido with routing rules that match severity and components.

### How to use

- Start with one catch-all, then split by severity as volume grows.
- Notify one for schedules; notify all for team-wide steps.
- Name policies exactly as referenced in schedules and runbooks.

### Policy design

# Escalation policy: [Name]

> Acknowledgement stops the chain. Unacked pages advance after the wait.

## Policy card
| Field | Value |
|-------|-------|
| Policy key | [stable-key] |
| Owner | [Team] |
| Reviewed | YYYY-MM-DD |

## Auto-routing (optional)
**Triggers:** Incident created · Incident updated

| Match | Value |
|-------|-------|
| Severity | Critical, High |
| Components | [API, Payments] |
| Tags | production |
| Teams | [Platform] |

*Leave empty for a catch-all. Split policies by severity to reduce alert fatigue.*

## Steps
| # | Target | Strategy | Wait | Notes |
|---|--------|----------|------|-------|
| 1 | Schedule: [Primary] | Notify one | 5 min | Current on-call |
| 2 | Schedule: [Secondary] | Notify one | 5 min | Backup |
| 3 | Team: [Platform] | Notify all | 10 min | Any responder |
| 4 | User: [Director] | Notify one | — | Final human |

**Targets:** On-call schedule · Team · User

**Strategy:** Notify one = round-robin · Notify all = parallel (first ack wins)

---

## Reference designs

**Production Critical**
Primary schedule (5m) → Secondary schedule (5m) → Team notify-all (10m)

**Business-hours Medium**
Match: Medium + tag `business-hours` → Primary schedule (15m) → stop

## Mid-incident
Use **Bring in** to add a user, team, or schedule without editing this policy.

## Test before prod
- [ ] Staging incident triggers correct policy
- [ ] Shift verification on for schedules
- [ ] Critical matches here; Low does not
- [ ] Waits feel right with real ack flow

## Avoid
Single step, no timeout · Same policy for every severity · Final step is one person with no backup


---

## 7. Severity matrix {#severity-levels-template}

Severity ends debates during outages. Align engineering, support, and leadership once — then wire the same levels into Incido and your escalation policies.

### How to use

- Review with support and product — severity should match customer experience.
- Wire levels into escalation routing (Critical ≠ Medium paging).
- Revisit after major incidents.

### Severity matrix

# Severity levels

| | Critical | High | Medium | Low |
|---|----------|------|--------|-----|
| **Customer impact** | Core workflow down for many | Major degradation, large subset | Limited; workaround exists | None or internal only |
| **Page on-call** | Immediately | Immediately | Business hours; optional off-hours | No |
| **Status page** | Yes | If visible | If visible | No |
| **Subscriber email** | Yes | Yes | Optional | No |
| **First public update** | ≤ 15 min | ≤ 30 min | If visible | — |
| **Follow-up** | Timeline review | Recommended | Optional | — |

## Examples
| Level | Example |
|-------|---------|
| **Critical** | Production API down · Payments/auth failure · Active security breach |
| **High** | Checkout success < [X]% · Regional outage · Primary API error spike |
| **Medium** | Non-critical feature flaky · Internal tooling down · Single-tenant issue |
| **Low** | Cosmetic bug · Monitor flake · Staging only |

## Decision tree (30 seconds)
1. Paying customers blocked on a core workflow? → **Critical**
2. Core feature degraded for many? → **High**
3. Limited impact or workaround? → **Medium**
4. Otherwise → **Low**

## Incido setup
- [ ] Create matching severity levels in org settings
- [ ] Link Critical/High to escalation policies
- [ ] Train support to use the same definitions


---

## 8. On-call runbook {#runbook-template}

What to check at 3 a.m. before you escalate. One runbook per critical service or high-volume alert — linked from monitors and on-call docs.

### How to use

- One runbook per noisy alert or tier-1 service — not one mega-doc.
- Test during game days; stale commands waste minutes.
- Link from monitor runbook field or team wiki.

### Service runbook

# Runbook: [Service / alert]

## Service card
| Field | Value |
|-------|-------|
| Owner | [Team] |
| On-call schedule | [Schedule name] |
| Escalation policy | [Policy name] |
| Status page component | [Component name] |
| Monitors | [Links] |
| Last tested | YYYY-MM-DD |

## What it does
[1–2 sentences — why a page here matters]

## If you see this
- [Alert condition or symptom]
- [User-visible behavior]

## Severity hint
| Condition | Level |
|-----------|-------|
| [All users down] | Critical |
| [Elevated errors > X%] | High |
| Otherwise | Medium — see severity matrix |

## Diagnose (in order)
1. [Dashboard/monitor] — check [metric]
2. Recent deploys — [CI/release link]
3. Dependencies — [DB, queue, vendor status]
4. Logs — `[query]`

## Mitigate
| Step | Action | Safe? | Expected result |
|------|--------|-------|-----------------|
| 1 | [Restart workers / scale] | Yes | Errors drop |
| 2 | [Rollback / feature flag] | Needs approval | Service recovers |
| 3 | Escalate after [X] min | — | Bring in [expert/secondary] |

## Rollback
[Link or exact steps · who approves]

## If customer-visible
Publish to **[component]** on status page:
> We are investigating [symptom] affecting [component].

## After resolution
- [ ] Timeline notes in incident
- [ ] Follow-up for missing monitors/docs
- [ ] Update this runbook

## Known false positives
- [Looks like outage but is not — how to tell]


---

## About Incido

Incido brings on-call schedules, escalation policies, structured incident workflows, and branded status pages together in one platform.

Learn more: https://incido.app
