What Is a Runbook? A Six-Part Format You Can Copy

A runbook is a step-by-step page for one known failure: how to recognise it, prove it, fix it and check the fix worked. It is written for the person on call at 2 a.m., so every step is a command they can copy, and every command has been run.
A runbook is not a wiki page about a system. It is not a list of everything that could go wrong. It covers one problem, the way it actually shows up.
The format here is the one used in the House of Loops fix-it runbooks: ten runbooks for a self-hosted n8n stack, each reproduced in a lab with its evidence log kept beside it. The worked example was reproduced again for this post on 25 September 2026, on n8n 2.40.5 behind Caddy 2.11.4, Docker 29.8.0 and Compose 5.5.1.
What is the difference between a runbook and a playbook?
Runbook vs playbook comes down to scope. A runbook answers "how do I fix this one thing?" A playbook answers "how does our team respond to this kind of incident?": who decides, who gets told, when to escalate, and which runbooks to reach for.
incident.io's guide (1 June 2023) draws the same line. A runbook "focuses on providing step-by-step procedures for resolving specific incidents in IT environments", while "a playbook is a broader strategic document that outlines an organization's overall approach to handling various situations."
For a small team, start with runbooks. A playbook with no runbooks under it is a meeting. Runbooks with no playbook still get your service back.
What goes in a runbook? The six parts
| Part | What it answers | Rule |
|---|---|---|
| Symptom | What does the person on call see? | Their words and the exact error text, not the cause |
| Confirm it | Is it really this problem? | Commands with the output to expect, and a table from result to fix |
| Cause | Why does it happen? | Short. Link the docs line that proves it |
| Fix | What do I change? | One change at a time, then straight to Verify |
| Verify | Did it work? | A check that fails when the fix did not work |
| Prevent | How do we not get here again? | The change to the setup, not the reminder to be careful |
Above the six parts sits an ingredients box: the stack, hardware, software versions, access, skills assumed, time to confirm and fix, and what it costs. The person on call can see at a glance whether they can do this now, or need someone with more access.
Two rules make the format work.
- The symptom comes first, in the reader's words. Nobody on call searches for "N8N_WEBHOOK_URL unset". They search for "webhook URL shows localhost".
- Verify must be able to fail. "The container is healthy" is not a check. In the example below, n8n reported
Healthywhile every URL it handed out was wrong.
A runbook template you can copy
# NN · <the symptom, in the words of the person who sees it>
> **Ingredients**
>
> | Ingredient | What you need |
> | ------------------ | ---------------------------------------------------- |
> | **Stack** | Which stack or service this covers |
> | **Software** | Versions it was written and re-run against |
> | **Access** | Shell, files, credentials the fix needs |
> | **Skills assumed** | For example: editing YAML, running docker compose |
> | **Time** | Minutes to confirm, minutes to fix, and any downtime |
> | **Monthly cost** | Usually $0 |
## Symptom
- What the reader sees, with the exact error text.
## Confirm it
Commands, each with the output a healthy system prints.
| What you see | Cause | Fix |
| ------------------------ | ----- | --------- |
| Result of a confirm step | Why | A, B or C |
## Cause
Two or three short paragraphs. Link the documentation that proves it.
## Fix
### A. The common case
One change, the command to apply it, then go to Verify.
## Verify
Checks that fail if the fix did not work. Say what "good" looks like.
## Prevent
The change to your setup that stops this happening again.
## Sources
Docs, with the date you read them. The lab run, with the date and versions.
Keep one runbook per file, numbered, in the same git repository as the stack it covers. Then a change to the stack and the change to its runbook can land in the same commit.
A worked example: n8n webhook URLs show localhost behind a proxy
This is the common case from the House of Loops runbook 07, cut down to Symptom, Confirm, Cause, Fix and Verify. The full runbook has more confirm steps, three more fixes and the Prevent section.
Symptom
You put n8n behind a reverse proxy such as Caddy. The Webhook node shows a production URL like https://n8n.example.com:5678/webhook/... or http://localhost:5678/webhook/.... You paste it into Stripe, GitHub or a form tool, and the calls never arrive. OAuth sign-ins fail the same way, because the callback URL carries the same wrong address.
Confirm it
From the folder with your compose.yml, load your hostname from .env, then see what the running container was actually given:
N8N_HOST=$(grep -E '^N8N_HOST=' .env | cut -d= -f2 | tr -d "\"'\r")
docker compose exec -T n8n env | grep -E '^(N8N_HOST|N8N_PROTOCOL|N8N_PORT|N8N_WEBHOOK_URL|WEBHOOK_URL|N8N_EDITOR_BASE_URL|N8N_PROXY_HOPS)='
In the lab, with the line deleted, it printed this. N8N_WEBHOOK_URL is missing:
N8N_HOST=n8n.localhost
N8N_PROXY_HOPS=1
N8N_PROTOCOL=https
Then open any Webhook node in the editor and read its production URL. The lab reproduced exactly the symptom. The settings n8n handed to the editor, and the URL a published test webhook reported for itself, were:
"urlBaseWebhook":"https://n8n.localhost:5678/"
"oauth2":"https://n8n.localhost:5678/rest/oauth2-credential/callback"
"webhookUrl":"https://n8n.localhost:5678/webhook/hol-echo-headers"
docker compose up -d --wait n8n still reported the container Healthy. A health check proves the process answers, not that its URLs are right.
Cause
n8n builds its own URLs from N8N_PROTOCOL, N8N_HOST and N8N_PORT, which default to http, localhost and 5678. Behind a proxy, n8n listens on 5678 while the world reaches it on 443, so it cannot guess the public address. The n8n docs say to set it by hand with N8N_WEBHOOK_URL. With the host and protocol set but no N8N_WEBHOOK_URL, and N8N_PORT left at its default, you get your host with :5678 added, as above. With none of them set, you get http://localhost:5678/.
Fix
Under services:, n8n:, environment: in compose.yml, make sure these four lines are there:
N8N_HOST: ${N8N_HOST:?set N8N_HOST in .env}
N8N_PROTOCOL: https
N8N_WEBHOOK_URL: https://${N8N_HOST}/
N8N_PROXY_HOPS: ${N8N_PROXY_HOPS:-1}
N8N_HOST in .env must be the bare name, with no https:// and no trailing slash. N8N_PROXY_HOPS is the number of proxies in front of n8n: 1 for Caddy alone, which sends the X-Forwarded-For, X-Forwarded-Host and X-Forwarded-Proto headers n8n needs by default. Add one for each further proxy, such as a CDN or tunnel. Then recreate n8n, the only process that reads these values:
docker compose up -d --force-recreate n8n
Do not set N8N_PORT=443 to make :5678 go away. n8n would then try to listen on 443 inside the container, and the proxy, which forwards to port 5678, would lose it.
Verify
Wait for n8n, then call a published production webhook of your own, in the same shell as the Confirm step so N8N_HOST is still set. Replace your-webhook-path with its path:
docker compose up -d --wait n8n
curl -sS -o /dev/null -w '%{http_code}\n' "https://$N8N_HOST/webhook/your-webhook-path"
You want a 2xx status, 200 unless your workflow is set to respond with something else, and the Webhook node's production URL should start with https://<your host>/webhook/ with no :5678. In the lab, after the fix, the same test webhook reported:
"urlBaseWebhook":"https://n8n.localhost/"
"oauth2":"https://n8n.localhost/rest/oauth2-credential/callback"
"webhookUrl":"https://n8n.localhost/webhook/hol-echo-headers"
The call returned 200. The n8n log had zero WEBHOOK_URL -> Use N8N_WEBHOOK_URL deprecation lines and zero ERR_ERL_UNEXPECTED_X_FORWARDED_FOR lines, and the lab's compose.yml was byte-for-byte the kit's file again.
That is the whole pattern. The symptom in the reader's words, a confirm step with the output to expect, one cause, one change, and a check that would have failed before the fix.
How do you keep runbooks honest?
- Reproduce the failure before you write the fix. Break it on purpose in a lab. If you cannot make it happen, you do not know the cause yet.
- Run the runbook as written. Copy each command from the page, not from memory. Every edit you needed is a bug in the page.
- Keep the output. A log file with the date and versions is what lets you say the runbook was tested.
- Re-run it when versions move. Runbook 07 was written on n8n 2.38.1 and re-run on 2.40.5 when the kit moved to that pin.
- Date every source. Docs change. "Checked on 23 September 2026" tells the next reader what to re-check.
If you run n8n yourself, the LM Studio vs Ollama post and the analytics post follow the same habit: every number comes from a dated lab run.
Frequently asked questions
What is a runbook?
A step-by-step page for one known failure: how to recognise it, prove it, fix it and check the fix worked, written so someone tired at 2 a.m. can follow it.
What is the difference between a runbook and a playbook?
A runbook fixes one specific problem. A playbook covers how a team responds to a class of incidents: who decides, who is told and which runbooks to use.
What should a runbook include?
Symptom, confirm it, cause, fix, verify and prevent, under an ingredients box listing the stack, access and time needed.
How do you know a runbook works?
Break the thing in a lab, follow the page exactly, and keep the output. Re-run it when the software changes version.
Why do n8n webhook URLs show localhost behind a reverse proxy?
n8n builds URLs from N8N_PROTOCOL, N8N_HOST and N8N_PORT, which default to http, localhost and 5678. Set N8N_WEBHOOK_URL to your public URL and recreate the container.
Sources
- Lab run, 25 September 2026: a scratch copy of the House of Loops
02-httpsstack (n8n 2.40.5, Caddy 2.11.4),N8N_WEBHOOK_URLremoved, n8n recreated, the fix applied, verified and torn down withdocker compose down -v. Log kept with this post's notes. - House of Loops fix-it runbook 07, "Webhook URLs show localhost or http behind the proxy", and its lab evidence on n8n 2.38.1 and 2.40.5.
- n8n docs, endpoints environment variables and webhook URLs with a reverse proxy, read 23 September 2026 for runbook 07.
- incident.io, What are runbooks and how do they fit into the incident management picture?, 1 June 2023.
The fix-it runbooks are a Premium course in the House of Loops classroom: ten runbooks for a self-hosted n8n stack, from disks that fill up to Code nodes that break after an upgrade, with the other fixes and the Prevent steps this page leaves out. See what the classroom covers on the syllabus, then join House of Loops free on Skool to start with the free courses.
Shannon Atkinson
House of Loops is a free community for people who would rather own their automation stack than rent it: n8n, Claude Code, AI agents, local models and the self-hosting underneath them, across 44 courses in the classroom.
Join Our Community

