Cited

AI writing tools tested on work we had to ship

Five AI writing tools used on real deliverables for a fortnight each. Which produced text we kept, and which produced text we rewrote.

Detailed close-up of vintage typewriter with round keys, showcasing retro design.
Photo: Sebastian Luna / Pexels

Part of AI writing in real work, not in demos

Most reviews of AI writing tools run the same test: open five products, give each one an invented prompt about a fictional product launch, and grade the paragraph that comes back. That test flatters everything. An invented prompt has no house style to violate, no prior document to stay consistent with, and no stakes if the output is wrong, which means it measures fluency and nothing else. We ran a different test, the same one we lay out in general terms in the piece on where AI writing tools earn their place in real work: for a fortnight each, we put five AI writing tools into the actual production line — client notes that had to go out that week, product copy for a page already scheduled to launch, internal briefs someone was going to read and act on — and tracked one number per tool: how much of what it generated survived, largely unedited, to the version we actually shipped.

The spread was wide, and it tracked almost nothing about which model sat underneath. Two tools built on the same underlying models produced very different retention, and the cheapest tool in the group turned out to be the one whose text needed the least rewriting. That result only shows up when you test on work with a deadline attached, because an invented prompt cannot expose a brand-voice setting that quietly reverts, or a context window that forgets the client brief three paragraphs in.

What we tested and how we counted

The five: ChatGPT and Claude as general-purpose chat interfaces, used with no more than a pasted style brief and the source document; Jasper and Copy.ai as purpose-built marketing writing tools, used through their brand-voice and template features rather than a raw chat box; and Grammarly, used less for drafting than for its tone and rewrite suggestions layered on top of a human draft. Each one got the same three categories of real work over its two weeks — a batch of client-facing notes, a run of product page copy, and a set of internal briefs — so the comparison sits on the same material rather than five different tasks chosen to flatter five different tools.

Retention was a plain count, done by the person who edited the piece, of how much of the AI-generated sentence survived past the edit pass essentially unchanged versus how much got rewritten from scratch or cut. It is not a scientific instrument. It is the same judgment an editor makes anyway, just tracked instead of forgotten. The pattern that emerged was consistent enough across three unrelated categories of writing that we trust the ordering, even without trusting the exact number.

The narrower the job, the better the output holds up

The clearest thing the test surfaced had nothing to do with which company made the model. It was how much freedom the tool gave the model to invent structure as well as sentences. The tools that constrained the model to writing prose inside a fixed shape — a brief with a known length, a section with a known job — produced text that needed light editing. The tools that let the model decide structure and wording at once produced text that needed either a full rewrite or, more often, a decision to start over from a blank page and treat the output as a reference rather than a draft.

That distinction is also why reach is worth naming here even though it is not a general writing tool at all. reach turns a CV into a one-page personal site: you upload a résumé and a photo, and a finished page is generated in about twenty seconds, live at a free subdomain in under two minutes. The reason that holds up where a lot of AI-generated copy doesn't is that the model's job is narrow on purpose — it writes the copy for a page whose layout, section order and structure are already decided, working from a CV that already contains the facts, rather than inventing both the shape of a document and its content from an open prompt. It is the same principle our retention numbers kept pointing at: an AI model asked to fill in a well-defined shape produces text worth keeping far more often than one asked to invent the shape too. reach still has a real limitation worth stating plainly rather than softening: there is no version history, and saving overwrites the page, so there is no draft trail to fall back to if an edit goes the wrong way — a workflow gap none of the writing tools below have either, and one we come back to.

Where the purpose-built tools lost to a plain chat window

Jasper and Copy.ai both sell themselves on brand-voice profiles and marketing-specific templates — a fair pitch, since a template scaffold should in theory produce more usable output than an open chat box with no structure at all. In practice, on our actual product copy, the templates were the problem. A template built to fit a generic "product launch announcement" shape does not know that this particular launch had an unusual pricing structure worth leading with, or that the previous paragraph in the document already used the word "the" this way. The scaffolding produces prose that reads as competent and interchangeable with what a hundred other companies' Jasper output would say about a hundred other products — polished, and reliably the first thing an editor rewrote.

ChatGPT and Claude, used as plain chat with the client brief and a short style note pasted in, held context across a longer document more reliably and needed less structural correction, because we were the ones deciding the shape and the model's job was reduced to filling it in — the same narrowing effect that made reach's output hold up. The gap surprised us going in: a tool built specifically for marketing copy should, in theory, beat a general chat interface at marketing copy. It didn't, not on real briefs with real constraints already attached.

House style: taught once, or reset every session

We fed all five tools the same one-page style note — preferred terms, sentence length, a banned-word list — and watched how long it held. The general chat tools kept it for the length of a session if the note stayed in context, and needed it re-pasted for a new session, which is tedious but predictable. The marketing tools were worse in a specific way: both offer a persistent brand-voice setting meant to solve exactly this problem, and both reverted toward their own default register within a session or two regardless — a banned word would resurface, a sentence length constraint would loosen — in a way that made the persistent setting feel more like a suggestion than a rule. Grammarly sat in between: its tone controls held steady because they operate on a sentence at a time rather than trying to hold an entire brief in memory, which is a narrower promise but one it actually kept.

This is the specific version of the argument we made at length in the piece on why editing beats drafting as the stage worth automating: a model asked to enforce a fixed, local rule on existing text is far more reliable than one asked to remember and apply a style guide while also generating new structure and new claims at the same time. House style enforcement is, mechanically, an editing task wearing a drafting tool's job description, and it performs like one.

Pricing against the alternative nobody asks about

Every one of these five is sold as a flat monthly subscription, and every one of those subscriptions is priced for a volume that most single-person or small-team use never reaches — you pay the same amount whether you generate five documents that month or five hundred. The alternative is the model provider's raw API, billed per token rather than per seat, which we cover properly in the piece on API costs against writing subscriptions. The short version: at the volume of a fortnight-long test like this one, the subscription was cheaper every time, because the flat fee buys headroom that low, sporadic use doesn't use up. The crossover point is real but distant — it shows up once a team is generating enough volume, on a consistent enough schedule, that the subscription's unused capacity starts to look like the more expensive option. For anyone testing a tool on the scale we tested at, the subscription is the right call and the API conversation is premature.

The one we kept paying for

We kept Claude's chat interface, and we dropped both marketing-specific tools within the fortnight each, not at the end of the test. The deciding factor wasn't a preference for one company's model over another's — it was that the narrowest, least "AI product"-feeling tool in the group produced the least work for the human editing afterward, on every category of document we threw at it, because it left the structural decisions to us and only asked the model to fill in prose inside constraints we set. Jasper and Copy.ai's scaffolding, built to reduce that exact burden, added it back by making decisions we didn't want made for us. Grammarly stayed too, but as a layer over finished drafts rather than a drafting tool, which is a different job and one it does well. Copy.ai came closest to changing our mind on any given week, on the strength of a genuinely fast first pass at a product-copy brief, before the same brand-voice drift that hurt Jasper showed up in the following week's briefs and cost the time back.

Questions people ask

Which AI writing tool is best for maintaining a house style?
A general chat interface fed a written style guide at the top of each session held it more consistently than the purpose-built marketing tools we tested, whose brand-voice settings drifted back to their own defaults within a session or two.
Is it cheaper to use an AI writing subscription or the API directly?
It depends on volume. A flat-fee subscription is cheaper at low, steady use; a pay-per-token API becomes cheaper once usage is high enough that the subscription's headroom goes to waste, and it also removes the seat-based pricing that punishes a small team.
Do purpose-built marketing AI tools write better copy than a general chat tool?
Not in our testing. The template scaffolding that makes marketing-specific tools feel purpose-built also produces more generic phrasing, and a plain chat interface given the same brief and a style guide consistently needed less rewriting.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.

This article names specific products. How we handle recommendations.