The n8n workflow reliability checklist
Eighteen checks for the failure that doesn’t turn a run red. Free, no signup, no email, nothing gated. Work through it on your own workflow in an afternoon.
How to use it
Take one workflow — the one whose failure would cost you most, not the most complicated one. Open the canvas and go through each check. Each has what to look for, why it matters, and what a bad answer looks like.
Score nothing. A total out of eighteen would be a made-up number, and the point is not a grade — it is finding the two or three that make you uncomfortable.
A. The failure that looks like success
The whole reason this list exists. n8n marks an execution successful when no node threw. That is not the same as the work having happened.
1. Does every path end somewhere that records what happened?
Look for: nodes with no outgoing connection that only shape data — a Set, an If branch that quietly terminates, a NoOp.
Why: when execution reaches one, nothing outside n8n’s own execution list knows the path was taken. If that branch is your “skip” or “duplicate” case, you have no count of how often it fires.
Bad answer: “It just ends there, it’s fine.” Fine until someone asks how many were skipped last month.
2. Can a node succeed and still produce nothing?
Look for: an HTTP Request whose response shape you don’t validate, followed by a mapping that reads fields off it.
Why: a clean 200 with a changed payload shape makes the mapping produce empty values. Nothing throws. The run is green and the record is blank.
Bad answer: “The API doesn’t change.” It changed for everyone else too.
3. Is there anything that asserts the expected shape, at the edge?
Look for: a validation step immediately after data enters — required fields present, arrays non-empty, timestamps within a sane window.
Why: validating at the edge turns a silent wrong result into a loud failure you can route. Validating nowhere means the first thing that notices is a human, later.
Bad answer: validation exists but only logs, and the item continues down the happy path anyway.
B. Error handling
4. Does any node that reaches outside n8n have an error output wired?
Look for: HTTP Request, database, email, any third-party
node. Check its error output, and its onError setting.
Why: with nothing wired and stopWorkflow set,
a failure halts the execution where it stands and the item in flight is not
written anywhere. It is not retried, not queued, not recorded. It is gone.
Bad answer: “It’ll show up in the execution list as failed.” Only if somebody reads the execution list.
5. Is there an error workflow configured — and has it ever fired?
Look for: workflow settings → Error Workflow, or an Error Trigger node somewhere.
Why: a global error workflow is the difference between a failure you learn about and one you discover in a customer complaint.
Bad answer: it is configured and has never fired. Either nothing has failed in months, or it is not wired the way you think. Test it deliberately.
6. Does a failure notify a person, or a channel nobody reads?
Why: alerting into a channel with three hundred unread messages is the same as no alerting, with extra steps and a false sense of coverage.
Bad answer: “It posts to #alerts.” When did someone last act on one?
C. Duplicates and idempotency
The most expensive category, because the damage is often financial and always invisible at the time.
7. Can the same input arrive twice?
Look for: webhooks from providers with at-least-once delivery — payment processors, e-commerce platforms, most webhook senders. Many retry if your response is slow, even when you did receive it.
Why: if the answer is yes and nothing deduplicates, the second delivery does the work again. Twice-charged, twice-emailed, twice-created.
Bad answer: “They only send it once.” Check their documentation rather than their behaviour so far.
8. If you deduplicate, what stores the key — and is the write atomic?
Look for: a Code node calling
$getWorkflowStaticData() that both reads and writes.
Why, and this is the one that surprises people: n8n gives each execution an in-memory copy of static data and writes the whole blob back after the execution ends. That is a read-modify-write race, and no code inside the node can close it, because the node never performs the write. Two concurrent executions can both pass the same duplicate check.
Bad answer: “We tested it and it caught the duplicate.” Sequential tests cannot establish a concurrency property. I know because that is exactly the mistake I made in my own workflow — the whole story is here.
9. Is the check-and-set a single operation?
Look for: a pattern of “does this exist? → no → create it” across two separate nodes.
Why: two executions can both read “no” before either writes. The fix is a store that decides atomically — a unique constraint, or an insert that does nothing on conflict — so exactly one execution wins by construction rather than by timing.
Bad answer: a delay or a random jitter between the check and the write. That narrows the window; it does not close it.
10. Is a retry safe to run twice?
Look for: retry settings on a node that sends, charges or creates something, and where the retried node sits relative to that send.
Why: retrying a read is free. Retrying a send is a second send unless the far end is idempotent.
Bad answer: retries enabled everywhere by default, without asking what each one repeats.
D. Triggers and entry points
11. How many entry points does this workflow have?
Why: every trigger is an independent way in, so a guard placed on one path does not apply to the others. Two triggers means two paths to audit, not one.
12. If it is a webhook, is authentication configured?
Look for: the Webhook node’s authentication parameter. In an export it is often absent rather than set to “none” — absent is the default and it means unauthenticated.
Why: the URL contains a generated id, so the protection is that nobody guesses it. A URL is not a secret once it has been in a browser history, a log, a proxy, a bug report, or a third party’s integration settings. Anything that reaches the endpoint enters your workflow as if it were the real caller.
Bad answer: “Nobody knows the URL.” That is a statement about today.
Something in front of n8n — a gateway, a reverse proxy, an allow-list — may already authenticate this. Worth confirming rather than assuming, in either direction.
13. If the webhook promises a response, does a node produce one?
Look for: response mode set to “using Respond to Webhook node”, with no such node in the workflow.
Why: the caller waits until it times out — and a caller that times out usually retries, which lands you back in section C.
E. Things that quietly change behaviour
14. Are any nodes disabled?
Why: n8n passes items straight through a disabled node, so the step is skipped while the diagram still shows it. The canvas says one thing and the execution does another. A disabled validation or dedupe step is the most expensive version of this.
Bad answer: “That was disabled for testing.” When?
15. Is any node unreachable, or fed by nothing?
Look for: nodes with no incoming connection, and Merge nodes with an input nothing feeds.
Why: an unreachable node is either dead weight or a branch someone meant to wire. A Merge waiting on an input that never arrives either stalls the branch or silently emits less than you expect.
16. Is there pinned test data left in the workflow?
Why: pinned data overrides what the node would actually produce. A workflow can run perfectly in testing and process the same frozen record forever.
F. Credentials and what leaves the building
17. Is any secret typed into a parameter rather than a credential?
Look for: API keys in a URL query string, in a header value field, or as a literal inside a Code node.
Why: credentials are stored separately and are not included when you export a workflow. Anything typed into a parameter is in the export — and exports get emailed, pasted into chats, and committed to repositories.
Bad answer: “It’s only in a header.” A header value is a parameter.
Before you send an export anywhere, to me or to anyone: open it and search it. This is the one item on this list where the cost of being wrong is not measured in lost records.
18. Could you rebuild this if the instance disappeared tonight?
Look for: where the export lives, who else can produce it, and whether anyone but the original author understands the workflow.
Why: most automation risk is not technical. It is that one person knew how it worked and no longer works here.
What this checklist cannot do
Stated as plainly as the rest, because a checklist that implies completeness is doing the same thing as a green run that did nothing.
- It cannot tell you whether any of this has actually happened to you. Every item above identifies a reachable failure. Only your execution history shows whether it fired, and an exported workflow does not contain execution history.
- It does not evaluate your expressions. A wrong
={{ ... }}is invisible to inspection and needs a run. - It does not read your Code nodes for logic errors — only for structural patterns like the static-data one in item 8.
- It knows nothing about your instance. A global error workflow, execution timeouts, queue mode and concurrency settings all change the answers, and none is in a workflow export.
- It is not a security assessment. Item 12 and item 17 are two specific, checkable things — not a claim that your setup is secure if they pass.
If you work through this and find nothing, that is a real result and you should believe it. A checklist that always finds something is a sales document.
If you want a second pair of eyes
Most of this list you can do yourself in an afternoon, and I would rather you did. Where it gets harder is the part a checklist is bad at: deciding which findings actually matter for your business, and being certain you have not missed one on a workflow with forty nodes.
That is what I sell — a fixed-scope review of one workflow, where every finding states what it does not prove, and the list of what I did not check is as plain as the list of what I did.
Tell me which items made you uncomfortable → Email. No form, no signup.
And whatever you do, don’t send secrets. Not to me, not to anyone. See item 17 — check your export before it leaves your machine.