Runbook
Why it matters
The first time something breaks, one person works it out under stress. If nothing is written down, the next occurrence needs the same investigation, and only that person can do it. A runbook turns a one-off discovery into a repeatable fix. It also lets someone else take over: a colleague, a contractor or an AI agent with the right access can follow the steps without calling the original expert.
How to apply it
- Write it straight after the first incident, while the details are fresh.
- Name the trigger in the title so the right runbook is easy to find when it counts.
- Lay out checks and decisions as a sequence: if A, do this, otherwise do that.
- Include where to look, such as which dashboard, which log and who holds the login.
- Keep it short enough to follow while something is actively broken.
- After each use, fix any step that was unclear or missing.
What it is
A runbook answers one question: when this particular thing goes wrong, what is the right next step? It starts with a named trigger, for example "the payment webhook has stopped arriving" or "a customer cannot log in". It then lists what to check, in order, and what to do depending on what is found. The best ones read like a short branching checklist, not an essay.
A runbook differs from a Standard Operating Procedure (SOP). An SOP describes a routine task done regularly, such as onboarding a client. A runbook describes an interruption, which may happen twice a year.
Common mistakes
- Writing it as a reference manual that nobody can scan in a crisis.
- Storing it somewhere that is unreachable when the failing system is the one that holds the documents.
- Never testing it. A runbook that has not been followed once may have gaps nobody has seen.