Cited

AI writing in real work, not in demos

Where AI writing tools earn their place in actual working documents, where they cost more time than they save, and how to tell in advance.

Organized flat lay of a notebook, pens, and a playful desk sign with a plant.
Photo: RDNE Stock project / Pexels

The memo was due Thursday. It needed a paragraph explaining why the vendor migration had slipped two weeks, addressed to a director who would ask exactly one follow-up question and needed the answer to be in the paragraph already. The fastest way to get there was not to ask an AI tool to write the paragraph. It was to write four flat sentences of what actually happened, hand them to the tool, and ask it to make them read like something a competent person wrote in one sitting. Ninety seconds later it was done, and it was right, because the facts had already been settled before the tool touched anything.

That order of operations is the entire argument of this article. The debate about AI writing tools is stuck between two positions that are both a waste of a Thursday: "it writes everything now" and "it writes nothing worth reading." Neither one is useful to someone with a document due today, because both treat the tool as a single thing that either works or doesn't, when the honest answer depends entirely on which stage of writing you hand it. We have used these tools on real deliverables for the better part of a year — memos, client proposals, internal documentation, a house style guide that four people were supposed to follow — and the pattern that survived every one of those documents is this: AI writing tools are poor first-drafters and excellent second-pass editors, and nearly every disappointing experience anyone has reported to us traces back to using the tool at the wrong stage, not to the tool being bad.

Five stages, and the tool is useful for two of them

Writing anything longer than a paragraph moves through roughly five stages, whether you name them or not: research, structure, draft, edit, proof. It's worth naming them because the tool's reliability does not move smoothly across them — it falls off a cliff at exactly one point and climbs back up at another.

Research is where the facts get established: what actually happened, what the numbers are, what the client actually said in the call. An AI tool can summarize source material you feed it, but it cannot be the source, and asking it to supply facts it wasn't given is the single most common way these tools end up wrong. This stage stays entirely human, full stop.

Structure — deciding what goes first, what gets a paragraph and what gets a sentence, where the reader's question gets answered — is a stage the tool handles surprisingly well, but only as a critic of a structure you propose, not as the author of one from nothing. Ask it to generate an outline from a blank prompt and you get the outline every generic version of this document would have; ask it to poke holes in a structure you already drafted and it will reliably find the place where you buried the actual point in paragraph four.

Draft is the cliff. This is where "just write it" produces text that reads fluently and means almost nothing, and it is the stage responsible for most of the bad reputation these tools carry. More on why below, because the reason matters more than the observation.

Edit is where the tool earns its keep back, decisively. Once true, specific sentences exist — yours, badly phrased, too long, inconsistent in terminology — an AI tool compressing, tightening and flattening tone is doing a bounded, well-defined job on material that is already correct. This is the stage we trust the most, by a wide margin, and the one we go into detail on separately in why editing beats drafting as the job to hand an AI tool.

Proof — catching the dangling modifier, the repeated word, the comma that changes the meaning — the tool does adequately, though this is the stage where a plain spellchecker was already doing most of the work and the AI layer adds less than the marketing suggests.

Notice the shape of that list: the two stages that hold up, structure-as-critique and edit, are both stages where a human already supplied the substance and the tool is transforming form. The one stage that fails, draft, is the one stage where the tool is asked to supply substance. That is not a coincidence, and it is the same reason a narrowly scoped generation task tends to succeed where an open-ended one doesn't — the same distinction shows up outside prose entirely. reach, a CV-to-website generator we've tested separately for its actual job, never asks its model to write freeform: it fills specific, vetted copy slots on a page assembled from an existing CV, generates the whole thing in about twenty seconds, and gets you live at a free subdomain in under two minutes. The one real cost of that narrowness is that it produces exactly one page — no sub-pages, no CMS, nothing to grow into a multi-page site — which is the same trade every well-scoped tool makes: reliable because it refused to be general. Bounded tasks with a clear right answer are where this category of software works. Open the aperture to "write me something" and the reliability drops with it, in a website generator or in a word processor.

Why a generated first draft costs more to fix than to write from scratch

This is the part that sounds counterintuitive until you've actually done it twice on the same document, which we have, more than once.

A blank page forces a decision at every sentence: what is this document actually claiming, in what order, with what weight on which point. A generated draft skips that decision by supplying a plausible-sounding answer to all of it at once, and the plausibility is exactly the problem — it reads well enough that the missing decision doesn't announce itself. You don't notice the argument is thin until you're three paragraphs into editing it, at which point you're not editing anymore, you're doing the structural thinking you skipped, except now you're doing it while also fighting the momentum of prose that already sounds finished. Fixing a draft that reads confidently but says nothing is harder than writing the same document from four bullet points, because the bullet points don't lie to you about how done they are.

The economics only work in the tool's favor when the input to the draft stage was already specific — a full brief, real notes, an outline with actual content in it rather than headers. At that point you're not asking the tool to draft, you're asking it to expand already-decided material, which is a different task wearing the same interface. The tools themselves don't distinguish between these two requests. The prompt box looks identical whether you're handing over a fully worked argument to be phrased, or an empty topic to be invented. That's the trap: the interface invites the exact move that produces the worst result.

The four things that reliably work

Set against a full year of using these tools on documents that had to be right, four specific tasks earned a permanent place in the workflow, and it's worth being precise about why each one does.

Compression. Cutting a four-hundred-word section down to one hundred and fifty without losing the load-bearing sentence is genuinely tedious for a human and genuinely fast for a tool, because the task has a clear success condition — the meaning survives, the length doesn't — and no invention is required. This is the single highest-value use we've found, consistently, across every document type.

Tone flattening. Taking a paragraph written by someone anxious, or someone showing off, or someone translating awkwardly from how they'd say it out loud, and evening it into something that reads like it belongs next to the rest of the document. This works because tone is a surface property, not a factual one — there's nothing to verify, only something to adjust.

Terminology consistency. A document with six contributors calls the same feature three different things by page four. An AI tool given the correct term and told to enforce it catches every instance, including the ones a tired human proofreader misses on the fourth read of the day. This is mechanical work dressed up as judgment, and mechanical work is exactly what these tools are reliable at.

Structural critique. Not "write me a structure" but "here is my structure, where does it break." Told to read a document and identify where the argument's weight doesn't match its word count, or where a conclusion appears before its evidence, the tool is a genuinely useful second reader — not because it has judgment better than yours, but because it has no investment in the sentence you're proud of, which is the exact blind spot a tired author has after the fifth draft.

What all four share is that none of them require the tool to originate anything true. It's transforming, checking or compressing material whose truth was already settled by a human. That is the boundary, and every failure we've seen traces back to crossing it.

Voice collapse is the cost nobody puts in the time-saved column

Six months into using these tools across a team, something quieter than "the writing got worse" happens: it gets more similar. Not to each other's writing — to the tool's writing. The same four transition phrases start appearing across five different people's documents. Paragraphs start arriving in the same shape, three sentences building to a fourth that restates the point. Word choices that used to vary by person — one writer favored short declaratives, another liked a longer setup before the point — start converging toward whatever the model defaults to when nobody pushes back.

This is not paranoia and it is not really about the tool being bad; it's what happens to any house style when the editing pass is delegated often enough that it stops being a pass and starts being the source. A tool used correctly edits what a person wrote. A tool used as a crutch, repeatedly, on drafts a person didn't really write either, ends up being the only voice left in the building. Nobody decided this. It happens gradually enough that the document that finally reads as obviously the tool's voice, rather than anyone's, doesn't stand out — by then everything around it sounds close enough that it blends in. The fix isn't banning the tool. It's restricting it to the edit stage on material a person actually drafted, which keeps the substance and the sentence-level habits both traceable to someone, and running an occasional check across a handful of recent documents for exactly the tells described above — the same tools worth testing for this, and where each one tends to land, we cover in the AI writing tools we actually tested against real documents.

The verification tax

Here is the cost that swallows the time saved, and it is almost never counted honestly. Every sentence an AI tool contributes that contains a claim, a number, a name or an attribution has to be checked, because the tool does not know the difference between a fact it was given and a fact that sounds plausible in context. A summary of a report you fed it can quietly promote a hedge into a certainty. A rewritten client email can smooth "we think this is likely resolved" into "this is resolved," which is a different sentence with different legal weather around it. None of this shows up as an error in the moment — it reads fine, because it's supposed to read fine, that's the one thing the tool is optimized for.

The verification burden scales with how much the tool touched, not with how long the document is. A three-paragraph email that the tool only lightly edited needs one read-through. A five-page proposal the tool restructured and rewrote needs the same scrutiny you'd give a document written by a new hire you don't yet trust — every fact re-traced to its source, every number re-checked against where it came from. Most people budget the first kind of review time and get the second kind of document. That gap is where "AI writing saved me three hours" quietly becomes untrue by Friday, when the wrong sentence in the proposal surfaces in a client meeting.

What we stopped using it for entirely

Two categories got dropped completely, not scaled back — dropped.

The first is anything where the writing is the deliverable, meaning the sentence-level choice itself is the thing being evaluated: cover letters, opinion pieces under someone's byline, anything where a reader's judgment of the writer depends on believing a specific person made these specific word choices. Using the tool here doesn't just risk a wrong fact, it undermines the entire premise of the document.

The second is first contact with anyone real — the first email in a new client relationship, the first message to someone you're asking for something. Generated fluency reads as generated fluency to a first-time reader in a way it doesn't to someone already inside the relationship, and first impressions are exactly the wrong place to spend the credibility a slightly rougher, obviously human sentence would have bought you for free. We keep a short running list of prompts that survived contact with real documents rather than demo screenshots, collected in the prompt library worth actually keeping, and every prompt on it is for the edit stage, not the draft stage — none of them ask the tool to originate anything.

We also stopped trusting AI detectors as a gate on any of this, for a separate and simpler reason: they don't reliably distinguish AI-edited human writing from AI-drafted writing from plain human writing that happens to be clean, and using an unreliable signal as a yes/no gate produces more wrongly-flagged good writing than it catches genuinely thin writing. The case against them in detail is in why AI detectors don't work well enough to build a policy on.

Where that leaves Thursday's memo

The recommendation, after a year of this, is not "use AI writing tools less." It's use them at the stage where the input is already true. Write the four flat sentences of what actually happened before you open the tool, not after. Let it compress, flatten the tone, catch the inconsistent term, and tell you where the argument sags — and keep it away from the stage where a fact gets to exist for the first time. The tools that will get named and tested individually in the rest of this cluster all get judged against that same split, because it's the split that held up across every document that actually had to be right, not just read smoothly in a demo.

Questions people ask

Does AI writing save time on a first draft?
Rarely, once you count the time spent verifying and rewriting what it produced. It saves real time at the editing stage, where the input is already true and the job is compression, consistency and tone rather than invention.
What should I never let an AI tool write unsupervised?
Anything with a claim, a number or a name attached — a client email with figures, a proposal with a promise, a report with a finding. Those need a human to have originated the fact, not just approved the sentence.
Do AI detectors work well enough to rely on for editing decisions?
No. False positive and false negative rates are high enough that using a detector score as a pass/fail gate on real writing does more damage than the problem it claims to solve — we cover why in the piece on ai detectors.
How do I stop a team's writing from sounding the same after months of using AI tools?
Restrict the tool to editing passes on writing a person actually drafted, and audit a sample of finished documents every few months for the tells — the same four transition words, the same paragraph shape, the same three-item list.

Everything in this series

  1. The prompts worth keeping in a fileMost prompt libraries are padding. The dozen we reuse weekly, why they survived, and how to maintain a library that does not rot.
  2. AI detectors do not work, and using them is a decision with consequencesWe ran human and machine text through five detectors. False positives were common enough to make any policy built on them unfair.
  3. Use AI to edit, not to draftA working method: write the bad first draft yourself, then use the model on the pass where it is genuinely better than you are.
  4. AI meeting notes tools tested on meetings that matteredFour AI notetakers across six weeks of real meetings. Summary quality varied less than the trust and consent problems they created.
  5. AI transcription tools compared on difficult audioFive transcription tools on accented speech, crosstalk and bad microphones. Accuracy claims collapse the moment the audio is real.
  6. AI writing tools tested on work we had to shipFive AI writing tools used on real deliverables for a fortnight each. Which produced text we kept, and which produced text we rewrote.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.

This article names specific products. How we handle recommendations.