Automation that survives a month
Most automations work for a fortnight and then quietly stop. What actually breaks them, and how to build the few that keep running unattended.

The automation is not broken. That is the first thing worth saying, because it is what makes this whole category of failure so hard to catch. The automation is fine. It was fine last month and it is fine now. It simply has not run since the fourteenth, and the reason nobody noticed is that its job was to do something quietly so that nobody would have to think about it, and it succeeded at the second half of that.
We have built a lot of these, for ourselves and while testing platforms, and the pattern is consistent enough to be depressing. A workflow gets built in an afternoon, works beautifully for two or three weeks, and then stops. Sometimes it announces itself — an error email, a red badge in a dashboard. More often it does not, because the way it died did not generate an error. And in the weeks between the stop and the discovery, somebody has been operating on the assumption that the leads were being logged, the invoices were being filed, the Slack channel was current. That assumption is the actual damage. The lost records are recoverable. The month of decisions made on top of a table that stopped filling is not.
So the useful question about any automation is not whether it works. On the day you build it, it works; that is why you stopped building. The question is whether anyone will find out when it stops. Everything else in this article follows from that one, and the honest version of the answer is that the difference between a workflow that survives a year and one that dies in a fortnight has very little to do with which platform you chose and almost everything to do with error handling and input discipline.
The strongest argument against all of this is that it is not worth the effort
Before the checklist, the counterargument, because it is a good one and we half agree with it.
The whole appeal of no-code automation is that a thing which would have been a two-day scripting job becomes a twenty-minute drag-and-drop. If you then spend three hours wrapping it in logging, alerting, retry logic and a heartbeat check, you have thrown away the economics that made you reach for the tool in the first place. At that point a small cron job and forty lines of code — which you can test, version, and read a year later — is arguably the better artefact.
That objection is correct for a large class of automations, and the correct response to it is not to harden them. It is to not build them. A workflow that saves you four minutes a week and takes three hours to make trustworthy has a payback period measured in years, and in practice it will be obsolete before it gets there because the process it automates will change. We come back to this at the end, because knowing which side of the line you are on is more valuable than any technique in between.
But there is a second class, and it is the one that matters: the automation that other people's work now depends on. Once anyone besides you has stopped doing the manual version because the automation exists, the cost of silent failure is no longer measured in minutes saved. It is measured in how long the wrong belief persists. For those, the extra hour is not overhead. It is the entire point.
Four things kill automations, and only one of them is your fault
Across everything we have watched break, the causes fall into four buckets. They are worth naming individually because they fail in different ways and need different defences.
Schema drift is the slowest and the most common. The API on one end of your workflow
changes what it sends. A field that was a string becomes an object. An optional field starts
arriving empty for a subset of records. A date format changes, or a nested key moves one
level up in a new API version. Nothing errors — your mapping just starts writing empty
cells, or writing the literal text [object Object] into a column, and the automation
reports success every single time. This is the worst failure mode in the whole taxonomy
because it is invisible from the platform's side. The run is green. The data is garbage.
Auth expiry is the fastest and the most honest. A token expires, a password rotates, somebody revokes an app in a Google Workspace admin console, an OAuth grant hits a provider's periodic reauthorisation policy. The connection goes dead and the platform usually tells you, because a 401 is unambiguous and every vendor handles it. The problem is not detection, it is delivery: the notification goes to the email address of whoever built the automation, which is frequently someone who has since changed roles or left. We have seen more automations die from an unread error email than from any technical cause.
Rate limits are the sneakiest, because they are intermittent. Your workflow runs fine on a normal Tuesday and fails on the one day someone bulk-imports two thousand rows. If the platform retries with backoff, you get a slow day and no harm. If it does not, you get a partial run: three hundred records processed, seventeen hundred dropped, and a success flag on the ones that made it. Partial success is much harder to spot than total failure, and it is the reason batch operations deserve more paranoia than event-driven ones.
A human renames a field. Somebody tidies up the Airtable base. Somebody adds a required
field to the HubSpot form. Somebody renames a Google Sheet tab from "Leads" to "Leads 2026"
because it is January. This one is not a technical failure at all, and no amount of retry
logic prevents it. It is an organisational failure, and the only real defence is that the
person who might rename the field knows the automation exists. Which, overwhelmingly, they
do not, because the automation lives in one person's account and has no visible footprint
in the tool being automated. Naming your sheet tab Leads — automated, do not rename is a
crude fix and it works better than anything clever.
Notice that three of the four originate outside your workflow. You cannot prevent them. You can only decide how loudly they announce themselves, which is why the rest of this is about detection rather than prevention.
An error you cannot see is worse than a crash
Every automation platform has a run history and an error notification. Turning that on is the floor, not the ceiling, and the reason it is not enough is a distinction that people consistently miss: a failed run and an absent run are completely different events, and only one of them generates an error.
If your workflow triggers, calls an API, gets a 500, and dies, that is a failed run. The platform sees it, logs it, and emails you. Good. But if the trigger itself stops firing — the webhook subscription lapsed, the polling connection lost its auth, the folder it watches got renamed, the app you connected changed its event model — then there is no run. There is nothing in the history. There is no error, because nothing happened, and nothing happening is indistinguishable from a quiet week.
This is the single most useful idea in this article, so we will state it as a rule: error notifications tell you when your automation misbehaves; they never tell you when it stops existing. You need a separate mechanism for the second one, and it has to live outside the automation it is watching.
The simplest version is a heartbeat. Have the workflow write a timestamp somewhere on every successful run — a cell in a sheet, a field on a record, a single-row table. Then build a second, dead-simple scheduled workflow that reads that timestamp once a day and messages you if it is older than it should be. That second workflow can be five steps long. It has no dependencies on the APIs the first one uses, so the failures that kill the first one do not kill the watcher. If you build exactly one piece of reliability infrastructure, build that.
Second: send the alerts somewhere a human actually looks, and somewhere more than one human looks. A shared Slack or Teams channel beats a personal inbox, not because it is more reliable but because it is socially harder to ignore. An error message sitting unread in a channel two colleagues can see gets acted on. The same message in your inbox at 2am on a Saturday gets swiped away.
Third: make the alert legible. "Scenario 4 failed" tells you nothing at 9am on a Monday. "Invoice sync failed — Xero step, record INV-2291" tells you whether to care. Most platforms let you build a custom error handler that formats the message; use it. The mechanics of setting these up differ enough between platforms that we have given them their own piece on error handling in no-code automations.
Idempotency, without the computer science
The word is off-putting and the idea is simple: running the same step twice should leave the world in the same state as running it once.
This matters because duplicate triggers are normal, not exceptional. A webhook delivery is retried because your endpoint was slow to acknowledge. A polling trigger picks up the same row twice because someone edited it and the "last modified" field moved. You re-run a failed scenario from the run history to recover the records it missed, and it cheerfully re-processes the ones it already handled. Every one of these is a routine event, and every one of them produces a second execution of a step that was designed assuming it would run once.
If that step sends an email, someone gets two emails and is mildly annoyed. If it creates a CRM record, you now have a duplicate contact and every downstream count is wrong. If it issues a refund, you have a real problem.
The defence is unglamorous. Carry a stable identifier from the source through to the destination — the source record's ID, the webhook's event ID, the message ID, whatever the originating system considers unique. Write it into a field on the created record. Then, before creating anything, search the destination for that identifier and branch: found means update, not found means create. This is two extra steps on most platforms, and it converts "create a contact" into "make sure a contact exists", which is a fundamentally sturdier instruction.
Where you genuinely cannot search the destination — some APIs make lookup expensive, some webhooks give you nothing stable to key on — keep a deduplication table of your own. One sheet, one column of processed IDs, one lookup at the top of the workflow. It is ugly and it works.
The other half of input discipline is validating before you write. Check that the fields you depend on are present and the right shape, and route anything that fails the check into a "needs a human" list rather than pushing it through. This is what catches schema drift: the day the API starts sending an object where it used to send a string, your validation step fires, the record lands in the exceptions list, and you find out on day one rather than in the quarterly report.
Run history is a debugging tool, not a safety net
Every platform keeps some record of what ran. It is genuinely useful — you can open a failed execution, see the payload at each step, and understand what went wrong without reproducing it. Use it. But be clear about what it is not.
Retention is a plan feature. It varies by vendor, it varies by tier within a vendor, and it changes when pricing changes, so we are not going to quote a number that will be wrong by the time you read this — check your own plan's page, because the figure on your account is the only one that matters. What is stable across all of them is the shape of the problem: history is finite, the cheaper your plan the shorter it is, and it is exactly the resource you run out of in the situation where you need it most. An automation that failed silently five weeks ago is an automation whose failure has very likely already aged out of the log you would use to diagnose it.
Self-hosting changes the terms of this rather than solving it. Running n8n on your own infrastructure means retention is a storage decision rather than a billing tier, which is genuinely better — and it means execution data now accumulates in a database you are responsible for pruning, which is a new job you did not previously have. It is a different trade, not a free one. We have written up how the three main platforms differ on this and on much else in the comparison of Zapier, Make and n8n.
The practical consequence: if a run produces information you will need later, do not let the platform's history be the only copy. Append a line to a log sheet — timestamp, record ID, outcome. It costs one step, it lives as long as you do, and it is queryable in a way run histories generally are not. It is also the thing that lets you answer "when exactly did this stop working", which is the first question anyone asks and the hardest one to answer from a truncated log.
Build the boring parts first, because you will not add them later
The order in which you build determines whether the reliability work happens at all. Build the logic first and you will ship the moment it works, and the logging you intended to add becomes a task that never reaches the top of any list. This is not a discipline problem. It is that a working automation generates no pressure to improve it, right up until it fails.
So invert it. Start with a workflow that triggers and does nothing except write a line to a log: timestamp, the identifier of whatever it received, and the word "started". Confirm the trigger fires when you expect and does not fire when you do not — this alone catches the polling interval you misunderstood and the webhook that sends on every field change rather than on creation. Then add the error handler and the notification, and deliberately break the thing to check the alert arrives where you want it. Then add the deduplication check. Only then write the step that actually does the work.
Built in that order, the reliability layer costs maybe twenty minutes and is finished before you have anything to be excited about. Built in the opposite order it costs the same twenty minutes and never happens.
One more habit, which sounds trivial and is not: write down what the automation does, in a sentence, somewhere the next person will find it. Not in the workflow's name — in the tool being automated. A note at the top of the sheet, a description on the Airtable base, a pinned message in the channel it posts to. The failure mode this prevents is the fourth one on our list, the human who renames a field, and it is the only defence against it that has ever worked.
Some things should stay manual
The last piece of reliability engineering is refusing to build.
A workflow that runs a handful of times a month, on a process that is still changing, saving a couple of minutes each time, is not worth hardening — and an unhardened automation on a process anyone depends on is worse than no automation, because it replaces a task someone knows they have to do with a belief that it is being done. Ask what happens if this silently stops for a month. If the answer is "we would notice immediately", build freely and keep it simple. If the answer is "we would not notice, and the consequences compound", either build it properly or leave the human in the loop deliberately.
Two shapes reliably do not justify the effort. Anything requiring judgement on more than a small minority of records, because the exceptions route back to a person anyway and you have built a machine that mostly generates exceptions. And anything on a process still being figured out, because you will rebuild the workflow every time the process moves, and the rebuild costs more than the manual version ever did. In a surprising number of cases the right answer is a well-structured sheet and a recurring calendar reminder, which is an argument we make properly in the case for when a spreadsheet beats an automation.
If you are starting out, the thing we would actually recommend is narrower than a methodology. Pick one automation — the one whose silent failure would hurt most — and give it the full treatment: heartbeat, shared-channel alerting, dedup key, its own log. Leave everything else exactly as it is. One workflow you trust completely is worth more than nine you half-believe, because the nine impose a running tax of doubt: every number they produce needs checking, which is the work you automated away. Our starting picks for that first one are in the list of automations worth building first.
The measure to hold on to is not uptime, which you cannot see anyway. It is the time between a workflow stopping and you finding out. Get that under a day and the rest of this stops mattering very much.
Questions people ask
- Why do my Zaps and scenarios stop working after a few weeks?
- Almost always because something upstream changed — a renamed field, an expired token, a new rate limit — rather than because the automation itself broke. The platform usually reports it, but only inside a dashboard nobody opens.
- How do I get alerted when an automation fails?
- Every major platform can email or message you on error, but that only covers runs that start and fail. Add a second automation on a schedule that checks whether the first one has run recently, because a trigger that stops firing produces no error at all.
- What does idempotency mean in an automation?
- It means running the same step twice produces the same end state as running it once. In practice that means writing a unique key from the source record into the destination and checking for it before you create anything.
Everything in this series
- When a spreadsheet beats an automationNot every repetitive task deserves a workflow. The volume, stakes and stability thresholds below which automating costs you more.
- Twelve automations worth building before any of the clever onesThe unglamorous automations that repay the effort every week, with the trigger, the steps and the failure mode for each.
- AI agent steps in automation tools: what they are good forEvery automation platform added an agent step. Where a model in the middle of a workflow helps, and where it makes failure unpredictable.
- Self-hosting n8n: two weeks in, and the bill in hoursn8n on a small VPS for a fortnight. Where self-hosting paid off, and the maintenance work the pricing comparison never mentions.
- Error handling for people who build automations in a browserRetries, dead-letter paths, alerting and idempotency, explained without code, plus how to add them in Zapier, Make and n8n.
- Zapier vs Make vs n8n after a fortnight on eachThree automation platforms, the same four workflows, two weeks each. Where the pricing model decides the answer before the features do.
- Webhooks explained for people who do not write codeWhat a webhook is, why it beats polling, and how to set one up and debug it in a no-code automation without touching a terminal.