Error handling for people who build automations in a browser
Retries, dead-letter paths, alerting and idempotency, explained without code, plus how to add them in Zapier, Make and n8n.

Part of Automation that survives a month
Open the run history of most no-code automations that have been live for more than a couple of months and you will find the same shape: a clean run of green checkmarks, then somewhere in the middle a red one nobody replied to, then more green ones, because the workflow kept going and just skipped the record that failed. Nobody built this on purpose. Nobody sat down and decided that a failed run should vanish into a log that gets pruned after thirty days. It happens because the platform gives you a trigger, an action and a publish button, and error handling is not on that path — you have to go looking for it, and by the time you go looking, something has already been lost.
The fix is not complicated, which is the whole point of writing it down. It is four pieces that fit together — retry, fallback, park, alert — and adding them to an existing workflow takes about twenty minutes. That is the retrofit this piece walks through, plus exactly where each of Zapier, Make and n8n hides the controls, because "add error handling" is easy to say and surprisingly fiddly to actually click through the first time. The reasoning for why this is worth doing at all — which automations deserve it and which do not — is the argument we made in full in automation that survives a month; this is the mechanism, not the argument.
The four-part pattern
Retry with backoff. Most failures in a running automation are not permanent. An API returned a 500 because it was briefly overloaded, a webhook delivery timed out on a slow acknowledgement, a rate limit tripped for thirty seconds. If the platform retries automatically after a short wait, the large majority of these resolve themselves and never need a human. A retry that fires immediately, with no wait, mostly just fails again for the same reason the first attempt did — the wait is not decoration, it is the part that gives the upstream system time to recover.
A fallback path. Retry a fixed number of times and then stop retrying — every platform that supports configurable retries also lets you cap them, and you should. What happens after the cap is exhausted is a separate decision, and it needs to be a decision rather than a default. The fallback is usually simple: send a Slack message instead of the CRM update it couldn't complete, or write to a backup sheet instead of the primary one.
A dead-letter store. This is the piece almost everyone skips, and it is the one that turns a failure into something recoverable instead of something lost. When a record cannot be processed after retries are exhausted, write it — the full payload, not just an ID — into a place designed to hold it: a table, a sheet, an Airtable base. "Dead-letter" is a term borrowed from message queues, where undeliverable messages get parked rather than dropped, and the no-code version of it is the same idea with worse tooling and no less value.
A human alert. Not a notification setting buried in account preferences that emails the person who built the workflow eighteen months ago and has since changed teams. An alert that lands somewhere a person actually looks today, worded specifically enough that they know whether to act on it now or on Monday.
None of this requires writing code. It requires four extra steps in a workflow you already have, in roughly this order, and the rest of this piece is where to click for each one.
Where each platform hides the error branch
The three platforms differ enough here that "check your automation platform's docs" undersells how different the actual workflow is. We cover the wider differences between them — pricing, retention, self-hosting — in the full comparison of Zapier, Make and n8n; this is just the error-handling slice of that.
| Platform | Where the error branch lives | Retry behaviour |
|---|---|---|
| Zapier | No native error branch on a step. The practical fallback is a Path that checks whether the previous step returned data, or a Filter step that only lets valid records through, combined with a separate notification step. | Largely automatic on paid plans; not step-by-step configurable the way the other two are. |
| Make | Right-click any module and add an error handler directly on it — Resume, Rollback, Break, Ignore or Commit, chosen per module. | A dedicated Retry error handler with a configurable interval and a capped number of attempts. |
| n8n | A separate workflow with an Error Trigger node, wired to any workflow that should report into it. Individual nodes also carry a "Continue On Fail" setting. | Retry-on-fail is a per-node setting, with a wait time between attempts. |
The practical read: Make gives you the most granular control because the error handler attaches to the exact module that failed, so you can decide different behaviour for different steps in the same scenario. n8n's separate Error Trigger workflow is a slightly different shape — it centralises failure handling across every workflow that points to it, which is useful once you have more than two or three automations and do not want to rebuild the alert logic in each one. Zapier is the odd one out: without a native error branch, the fallback path has to be engineered from Filters and Paths rather than configured directly, which is more workaround than feature. If your workflows are simple and few, that gap does not matter much. If you are running a dozen Zaps that other people depend on, it is a real argument for Make or n8n, not just a footnote.
Why the same trigger fires twice, and what to do about it
Idempotency is the property that running a step twice leaves the world the same as running it once, and the reason it matters here is that duplicate triggers are routine, not rare. A webhook sender retries delivery because your endpoint acknowledged slowly — the sender assumes failure and sends again, and now you have two identical events for one real occurrence. We go into why that happens on the sending side in webhooks explained for no-code builders. A polling trigger picks up the same row twice because someone edited an unrelated field and the "last modified" timestamp moved. And the most common cause of all: you replay a failed run from history to recover the records it missed, and it cheerfully reprocesses the ones that already went through, because the workflow has no memory of what it already did.
The fix does not need code. Carry a stable identifier from the source through to the destination — the source record's ID, the webhook's event ID, whatever the origin system considers unique — and write it into a field on whatever you create. Before creating anything, search the destination for that identifier first: found means update, not found means create. That is two extra steps, and it converts "create a contact" into "make sure a contact exists," which behaves correctly no matter how many times it runs. Where the destination cannot be searched efficiently, keep your own dedupe table instead: one sheet, one column of processed IDs, checked at the top of the workflow before anything else happens.
The parking table, and how to replay from it
The dead-letter store does not need to be sophisticated. A single Airtable base or a Google Sheet with columns for timestamp, source record ID, the step that failed, the error message, and the full payload as JSON text is enough. The payload column is the one people skip and the one that matters most — without it, "this record failed" tells you a record exists somewhere that needs attention, but not what it contained, which means recovering it means going back to the original source rather than just replaying what you already captured.
Replay is the reason the table has to hold the whole payload, not just an ID. Once the underlying problem is fixed — the field is renamed back, the token is refreshed, the rate limit window has passed — you need a way to feed those parked rows back through the workflow without manually re-triggering each one. The cheap version of this is a second, small workflow: it reads unprocessed rows from the parking table, sends each payload through the same logic the main workflow uses, and marks the row done on success. It looks like duplicated work at first, but it is a workflow you will use rarely and be very glad to have when a schema change parks forty records at once instead of one.
The twenty-minute retrofit
For a workflow you already have running, in order:
Add the retry setting on the step most likely to fail — usually the one making an external API call — with a short wait and a small cap, two or three attempts. Add a fallback branch after the cap: on Make, an error handler on that module; on n8n, continue-on-fail plus a check downstream; on Zapier, a Filter or Path that catches what didn't complete. Point that fallback at a dead-letter row containing the full payload, the timestamp and the error text. Add one notification step that posts to a shared channel rather than an inbox, worded with the specific record or step that failed rather than the workflow's generic name. Then add the dedupe check at the top, before anything gets created.
That is five steps on top of a workflow that already works, and it is the difference between a failure you find out about from an angry client three weeks later and one you find out about from a Slack message the same afternoon, with the exact record already sitting in a table waiting to be replayed.
Questions people ask
- What is a dead-letter path in a no-code automation?
- A branch that catches records your workflow could not process and writes them somewhere durable — a row in a sheet or an Airtable record — instead of letting them silently disappear or crash the whole run.
- Do I need to add retries manually in Zapier, Make or n8n?
- Make and n8n let you configure retry behaviour per step or per node with a wait interval and a retry count. Zapier's retry logic is largely automatic on paid plans and is not user-configurable the same way, so the fallback path matters more there.
- What does idempotent mean for an automation trigger?
- It means the same trigger event can fire the workflow twice — which happens routinely with webhooks and polling triggers — without creating a duplicate result, because the workflow checks for an existing record before creating a new one.
- How long does it take to add basic error handling to an existing workflow?
- About twenty minutes for a workflow with a handful of steps — a dead-letter destination, one error branch, one alert, and one dedupe check. It does not require rebuilding the workflow.