AI agent steps in automation tools: what they are good for
Every automation platform added an agent step. Where a model in the middle of a workflow helps, and where it makes failure unpredictable.

Part of Automation that survives a month
We built the same three tasks in three platforms twice each: once as a chain of branches and lookups, once with an AI agent step doing the equivalent work. Then we ran both versions against the same eighty-odd real inputs — support tickets, invoices, a folder of scanned forms — and watched what came out the other end. The honest headline is that the answer is not "agents are the future of automation" or "agents are a fad." It depends entirely on what happens after the model produces its output, and specifically on whether that output is a label or an action.
The counterargument first, because it is mostly right
The case against putting a model in a workflow is simple and correct for a large share of what people build with one: a branch is auditable and an agent step is not. You can read an if/else chain, know exactly which condition sent a record down which path, and be certain that the same input takes the same path every time it runs. A model step gives you none of that for free. It reasons in a way you cannot fully inspect, and the same input can produce a different output on a Tuesday than it did on a Monday. For anyone who has spent a career being told to keep business logic deterministic and testable, adding a language model to the middle of a pipeline reads as giving up ground you already held.
That objection is not wrong, and we are not going to spend the rest of this piece talking you out of it. It is simply incomplete, because it treats "business logic" as one category when it is really two. Some of what automations do is genuinely rule-based: this vendor's invoices always have the total in the same position, this webhook always carries a status field with one of four values. For that category, a branch is not just safer than an agent step, it is also faster, cheaper, and easier for the next person to read — do not replace it. The other category is where the input itself resists rules: free text written by a customer, a scanned form with a layout that varies by year, a support ticket that could be about a refund, a bug, or a compliment depending on wording nobody standardized. That second category is where the three weeks of testing changed our minds.
Classification and extraction: the clear win
We tested three task shapes on purpose, because they fail differently and an agent step earns its place in exactly two of them.
Classification — reading a support ticket and routing it to billing, technical, or account access — is where the win was largest and least ambiguous. The hand-built version needed a growing list of keyword matches and regex patterns to catch phrasings nobody anticipated when the rules were written, and it kept missing the ones written in a hurry, with typos, or in a tone that didn't match the training examples. The agent version read the ticket and picked a category correctly on cases the keyword chain simply had no branch for — sarcasm, a complaint that mentioned billing only in passing, a ticket that used a product nickname the rule list didn't know. This is the case the "dozen brittle if-branches" mustCover point is really about: past a certain number of edge cases, a keyword list stops being a list of rules and becomes a list of things you noticed broke last week. A model replaces that list with something closer to actual reading comprehension, and for a low-stakes routing decision — worst case, a ticket needs one manual reassignment — that trade is a clear win.
Extraction — pulling a total, a date, and a vendor name out of an invoice whose layout is not fixed — told a similar story. Where the layout was consistent, a template-based parser did the job with no model involved and no reason to add one. Where it wasn't, because invoices arrived from forty different vendors in forty different formats, the agent step handled variation the template parser had no mechanism for at all. It is worth being honest about the failure mode here, because it exists: on a badly scanned or unusually laid-out invoice, the model would occasionally return a number with real confidence that was simply wrong — not a crash, not a blank field, a wrong answer that looked exactly like a right one. A rule-based parser fails loudly on the same document; it returns nothing and you notice. That difference is the entire argument for the next section.
Tool-calling with side effects — letting the agent decide to issue a refund, update a CRM record, or send an email on its own judgement — is the one we stopped testing in production conditions after the first week, and it deserves its own section rather than a paragraph here.
The same input, two different Tuesdays
The operational problem with a model step is not that it is sometimes wrong. Every step in an automation is sometimes wrong; that is what error handling in automation that survives a month is for. The problem is that a model step can be wrong inconsistently, on the identical input, without any change to your workflow, your data, or the world. We fed the same ticket back through the classification step nine days apart and got two different categories, with no version change on our side. Nothing was broken. Nothing errored. The model just weighed the same words slightly differently the second time.
For a branch, that sentence does not parse — an if/else on a fixed condition returns the same result on Monday and on the Tuesday nine days later, forever, by construction. That guarantee is not a nicety, it is most of what makes an automation debuggable: when something downstream looks wrong, you can reproduce the input and watch the same path fire again. An agent step breaks that guarantee quietly, and it does so in a way that looks identical to correct, non-deterministic behaviour right up until it routes a genuinely urgent ticket to the wrong queue on the one day it matters. This is not an argument against agent steps. It is an argument for treating "this step is non-deterministic" as a fact you design around rather than a detail you discover in a postmortem, in exactly the spirit of the error-handling piece on dead-letter paths and alerting.
Cost per run stops looking like a rounding error
A branch costs essentially nothing to execute — it is a comparison, not a computation. A model call costs tokens, and tokens are the one part of this that scales with the size of the problem rather than the number of runs. A short classification prompt against a short ticket is a trivial cost per run. The same agent step asked to also summarize the ticket, check it against three prior tickets for context, and draft a suggested reply is now three or four model calls deep for a single input, and that cost compounds across every run of a workflow that might fire hundreds of times a day. The platforms make this easy to miss, because the per-run cost sits inside a task or operation count alongside every other step, with nothing flagging that this particular step is the one driving the bill. The practical rule we landed on: instrument the agent step's cost separately from the rest of the workflow from day one, because by the time it shows up as a surprising total, you have already paid for the runs that caused it.
Constrain the output before you trust anything downstream
The single change that made agent steps usable in production was refusing to let the model return free text and instead forcing a schema: a fixed set of allowed category values for classification, a typed JSON object with required fields for extraction, nothing else accepted. Every platform we tested supports this in some form — a JSON schema on the output, an enum of allowed values, a validation step immediately after the model call that rejects anything outside the shape and routes it to a human queue rather than letting it flow downstream. This does two things at once. It catches the cases where the model hallucinates a field that doesn't exist or returns a category that isn't one of the real options, and it makes the step's failure mode look like every other step's failure mode — a validation error you can alert on — instead of a silent wrong answer that only announces itself when someone downstream notices the total is off.
What a schema does not do is fix non-determinism. A constrained output can still be a different constrained output on a different run. That is the limit worth being honest about: schemas make wrong answers legible, they do not make the model more consistent.
The rule for letting a model touch a production system
Everything above argues for a single, fairly narrow rule, and it is the one we now apply before wiring any agent step into a live workflow: a model may choose a label, a category, or a value that a human or a downstream rule then acts on, and it may never be the thing that fires the action with the side effect. Classification and extraction fit this rule naturally — the model reads, a subsequent branch decides what happens with what it read, and the branch is exactly as deterministic and auditable as it would have been without the model in the loop. Refunds, record deletions, and outbound emails do not fit it, because the cost of a wrong output there is not "reroute a ticket," it is "money left a real account because a model was confident about the wrong number on the one bad scan in a thousand." Put an approval step between the agent's suggestion and anything that touches money, a customer record, or an external system, even when that feels like it defeats the purpose of automating the task. It does defeat part of the purpose. It is also the difference between an agent step that saved you a dozen brittle branches and one that quietly did something you cannot explain to the person who has to fix it.
If you are running any of this self-hosted rather than on a vendor's cloud plan, the maintenance calculus changes again in ways worth reading before you commit — we logged what that actually costs in hours in self-hosting n8n: two weeks in.
Questions people ask
- Should I use an AI agent step or a branching if/else in my automation?
- Use branching wherever the categories are fixed and few. Reach for an agent step only when the input is unstructured text or the number of branches would be too large to maintain by hand — and constrain its output to a schema either way.
- Can an AI agent step in Zapier, Make or n8n take actions on its own?
- All three can be wired so the model's output triggers a real action — sending an email, updating a record, issuing a refund. Nothing in the platform stops you from doing this; the discipline has to come from you, in the form of a human approval step before anything with a side effect fires.
- Why did the same automation produce two different results on the same input?
- Because a model step is probabilistic, not deterministic — the same prompt on the same input can select a different branch or word a summary differently from one run to the next, which a hand-built if/else chain never does.
- Does an agent step in a workflow get more expensive as the workflow scales?
- Yes, and non-linearly if the step is chatty — every extra field you ask the model to reason about or every retry it triggers adds tokens, and that cost compounds across every run of a workflow that might otherwise cost a fraction of a cent.