A Runbook for When Your Brain Goes Offline
When a service goes down, the first thing that breaks isn’t the server. It’s your brain. Adrenaline hits. Tunnel vision sets in. You know you should do something, but you freeze.
I’ve seen smart engineers freeze. I’ve done it. A runbook tells you what to do, in what order, so you don’t have to think.
This is the runbook I built. It runs on OODA: Observe, Orient, Decide, Act. One reason: make the obvious automatic when nothing feels obvious.
Why it matters
Most teams don’t have a runbook, just a shared understanding of “what we usually do.” That works until 2 AM when the pager goes off and the handler is asleep.
A runbook does three things:
- Removes the decision tax. Execute the next step. Don’t decide what it is.
- Prevents the “I’ll investigate first” trap. An agent sits on a problem for 45 minutes, figuring it out alone.
- Creates a shared rhythm. Everyone knows what to do at T+5, T+10, and T+30.
The five questions
Before you declare an outage, answer five yes/no questions:
- Spread: one customer, or multiple unrelated customers?
- Impact: is a core feature broken?
- Timing: is it happening right now, and can you reproduce it?
- Confirmation: can you see it yourself, or do monitoring alerts confirm it?
- Workaround: is the customer fully blocked with no alternative?
If three or more point to YES: tell your team lead immediately. Don’t investigate alone. Don’t try to be the hero who fixes it before anyone notices.
This is the single most important rule. In the example below, following it shaves 46 minutes off the response time.
Severity: who decides
The agent who detects the issue reports facts. The team lead classifies severity. Never let the same person do both.
- Critical: system down, no workaround, all users affected
- Partial: major feature degraded, many users impacted, no reasonable workaround
- Incident: minor impact, some users, workaround available
Severity drives everything: notification cadence, escalation path, who gets pulled in. Getting it right matters. Getting it fast matters more.
How it works: the timecode rhythm
The runbook runs on a clock anchored to T+0, the moment of detection:
T+0 Detect + Report Agent gathers facts, tells team lead
T+5 Appoint + Escalate Lead assigns roles, tech lead escalates
T+10 Open War Room IC opens a dedicated chat channel
T+30 First Notification "We're investigating" goes to customers
↓
Every 30 min: Update customers — even with no new information
↓
Resolution → Monitoring notification (within 10 min)
Stable for 30–60 min → Resolved notification (within 10 min)
The rhythm matters more than any single message. Every 30 minutes, customers hear from you, even if only “still investigating.” Silence is the worst thing you can send.
Roles: who does what
Three roles, mapped to the OODA loop:
- Incident Commander (IC): the team lead. Owns the incident. Approves every communication. Runs the war room. One IC at all times — never two, never zero.
- Tech Lead: escalates to engineering, stays with them until resolved. The bridge between support and the people who can fix things.
- Scribe: drafts customer notifications. Flags incoming tickets. The IC approves; the scribe sends.
The IC isn’t the smartest person in the room. The IC is the person who makes sure the smart people are talking to each other.
The OODA loop maps directly to concrete steps:
- Observe (Steps 1–2): detect, gather facts, report. No judgment, just facts.
- Orient (Step 3): classify severity. This is where the team lead makes the call.
- Decide (Steps 4–5): announce the outage, appoint roles.
- Act (Steps 6–14): escalate, log tickets, open the war room, notify customers, update every 30 minutes, monitor, resolve, review.
A real example (generalized)
A media processing service went down. Customers across Europe couldn’t generate subtitles. First trouble ticket: 11:30 AM.
The agent acknowledged but didn’t escalate. At 12:01 the ticket was still being worked solo. Outage declared at 12:16 — 46 minutes after the first signal.
Once the runbook kicked in: roles assigned, engineering pulled in, war room opened. But T+0 delay cascaded. First customer notification: almost two hours late. Then 50 minutes of silence.
Apply the runbook from the start:
- The agent escalates at T+0 instead of investigating alone. 46 minutes saved.
- Customer comms go out within 30 minutes, not two hours.
- The 30-minute update cadence prevents the silent gap.
The service was restored. The retrospective was honest. And the runbook got written.
Common mistakes
- “I’ll figure it out first.” The most expensive four words in incident response. Report first, investigate second.
- Delayed customer communication. Every minute of silence erodes trust. Send “we’re investigating” before you know anything useful.
- No formal IC handover. When an outage crosses shifts, the clock doesn’t reset. Timecodes stay at the original T+0. Outgoing IC posts a brief; incoming IC acknowledges. One IC, no silent handoff.
- Updates only when there’s news. If 30 minutes pass with nothing new, send “still investigating, no updates yet.” The customer needs to know you haven’t forgotten them.
- Forgetting the runbook exists. Read it before the outage. Nobody reads documentation for the first time during one.
Key takeaways
- Three of five questions point to YES → declare. Don’t investigate alone.
- The agent reports facts. The lead decides severity. Never the same person.
- T+30: first customer notification. Update every 30 min after, even with no news.
- Shift handovers must be explicit: one IC, no clock reset, no silent transfer.
- A runbook makes the obvious automatic under stress. Read it before you need it.