Cited

Use AI to edit, not to draft

A working method: write the bad first draft yourself, then use the model on the pass where it is genuinely better than you are.

A woman in a suit uses a red pen to edit a printed document on a green table with a laptop.
Photo: cottonbro studio / Pexels

Part of AI writing in real work, not in demos

The instinct to hand a model a blank page and a topic is understandable, because a generated draft feels like progress in a way a blank cursor does not. Fifteen seconds after asking, you have four hundred words with headings and transitions and something that reads like an argument. We treated that as a head start for the better part of a year, on real documents with real deadlines, and the feeling isn't wrong — a generated draft is progress, on paper. What it is not is cheap progress, and the gap between those two things is the subject of this piece.

The same distinction shows up outside prose. reach, a CV-to-website generator, splits its own generation into two phases instead of asking one model call to do everything: a deterministic pass that assembles a complete page from vetted layout pieces with no AI involved, and a second pass — a single model call — that writes the copy into slots the first pass already built. The whole run, both phases combined, finishes in about twenty seconds. The model is never asked to invent a page from nothing; it fills a bounded space it didn't have to design. That's the split this article argues for in writing: narrow tasks are where a model is reliable, open-ended ones are where it merely sounds reliable. The narrowness has a real cost, worth naming rather than hiding — reach produces exactly one page, with no sub-pages and nothing to grow into.

Why editing a generated draft costs more than writing a bad one yourself

The mechanism is worth stating plainly, because "AI drafts are unreliable" undersells the problem. A model writing a first draft predicts a plausible continuation, sentence by sentence, weighted toward whatever is statistically common in similar documents. The output is fluent by construction, and indifferent by the same construction to whether any given claim is true — a wrong figure sits in the same confident register as a correct one, and nothing on the page tells you which is which.

Editing your own bad draft is fast because you have an internal map of where the weak spots are: you remember hesitating over the third paragraph, and that the number in the second section came from memory and needs checking. Editing a model's draft means auditing every sentence with equal suspicion, because that internal map doesn't exist — you have no memory of writing it. That audit is slower than writing from your own notes, not faster, on anything where being wrong costs more than being late. A placeholder outline you intend to gut is the exception: the value was never in the sentences, it was in having a structure to react against, and a model producing plausible junk to argue with is a legitimate use of a drafting tool.

We make the broader case, across all five stages of writing, in AI writing in real work, not in demos. This piece is narrower: a method, not a theory. Draft by hand. Use the model only on the pass after that.

The four passes worth delegating

Editing is not one job. Four passes hold up on every document we've run through this, regardless of subject, and what they share is that none require the model to introduce a new fact — all four reshape material you already know is true rather than originating it.

Compression. Cutting a paragraph to two-thirds its length without losing the claim inside it is mechanical, well-specified work — "shorter, same meaning" is the target, and well-specified targets are where a model is reliable rather than merely fluent.

Structural critique. Asking whether an argument holds together, or which section is weakest, produces answers worth taking seriously more often than not. This doesn't require the model to know anything true about the world; it requires recognizing whether a sequence of claims follows a familiar, convincing shape, which is close to what these tools are built to do.

Terminology consistency. Catching that a document calls the same thing three different names — a "workspace" on page one, an "account" on page four, a "project" in the conclusion — is pattern-matching a model handles better than a human on a fourth read-through, tired and blind to it from over-familiarity with their own draft.

Final proof. The narrowest and most misunderstood of the four. A model catches grammar, awkward repetition, and dangling references reliably. It does not catch a wrong date, because proofing for grammar and proofing for fact are different jobs, and a fluent sentence with a wrong number reads as clean to a tool checking fluency only. Use this pass for surface mechanics — the models and their failure rates on exactly this task are in the AI writing tools we actually tested.

The six prompts we reuse

These are the actual prompts, not paraphrased. We keep them in a shared document rather than retyping them — a prompt retyped from memory drifts a little each time, the same way an oral story does.

1. Compression. "Cut this paragraph to roughly two-thirds its current length. Do not remove any claim, number, or attribution — only remove words that aren't doing work. If a sentence can't be shortened without losing something, leave it as is and tell me which sentence that was."

2. Structural critique. "Read this draft as a skeptical editor, not a collaborator. Tell me which section is weakest and why, whether the order of sections serves the argument, and whether any claim appears without support. Don't suggest rewrites yet — just diagnose."

3. Terminology consistency. "List every term in this document that refers to the same underlying thing but is worded differently across sections. For each one, tell me where it appears and what the different wordings are. Don't pick a winner — I'll decide."

4. Final proof. "Check this for grammar, repeated words, and dangling references only. Do not comment on tone, argument, or word choice, and do not flag anything you're unsure is actually an error."

5. The voice guard, prepended to every one of the above. "Do not rewrite any sentence you weren't asked to change. Do not add transitions, summary sentences, or rhetorical questions. Do not introduce a three-item list where the original had a different number of items. If you're tempted to smooth something for rhythm, leave it rough instead."

6. The meaning-drift check. "Here is my original paragraph and your edited version, side by side. List every place where the edited version implies something the original didn't say, drops a qualifier the original had, or changes the relationship between two claims — even subtly. Ignore phrasing; only flag meaning."

Guarding your own voice

Prompt five exists because the cost that's easiest to miss doesn't show up in any single document — it shows up across six months of them. A model has a small number of default rhythms: the three-item list that exists for cadence rather than content, the sentence that opens with "here's the thing," the closing line that restates what the paragraph just argued. None of these ruin one piece. Run enough editing passes through the same tool without forbidding them, though, and a publication's voice drifts toward the tool's median output, a few hundred small nudges at a time, all pointing the same direction.

The fix is maintaining an explicit list of moves you no longer accept from an edit, stated in the prompt every time, updated whenever you catch a new tic. We keep ours next to the six prompts above rather than trusting memory. It's also why we stopped relying on AI detectors to spot machine-smoothed prose after the fact — the reasons are specific and mechanical, and we've written them up in why AI detectors don't actually work rather than repeating the case here.

The check that catches the model quietly changing your meaning

This is the step most people skip, and it's the one that matters most. Prompt six exists because an editing pass that's supposed to only tighten a sentence will sometimes, while tightening it, drop a qualifier ("usually" becomes nothing), flip an implied causation ("which suggests" becomes "which shows"), or quietly narrow a claim you meant to leave broad. None of this gets flagged by the model doing the editing — it isn't lying, it genuinely doesn't track the difference between rephrasing and re-arguing. The only reliable catch is a second pass, run separately, whose only job is comparing meaning between the two versions. We run this on anything longer than a paragraph before accepting an edit, and it catches a real change roughly as often as it finds nothing.

Where this method fails

Three cases, worth naming rather than glossing over, because a method with no acknowledged failure mode is one nobody should trust.

Anything highly technical, where confident-sounding compression can quietly flatten a distinction that mattered — a "may" that was load-bearing, a condition that applied to one case and not the general one. Compression optimizes for shorter and readable; it does not know which nuance was the point.

Anything legally load-bearing — contract language, terms of service, a regulatory disclosure. The burden isn't checking facts, it's checking that a rephrased clause still means what it did under the specific legal reading it was drafted for, which a general-purpose model can't judge.

And numbers you haven't verified before the pass begins. Every prompt above assumes the fact is already right; none substitute for getting it right first. An unverified number that reaches an editing prompt comes out reading more confident, not more correct.

None of this argues against the method — it's the boundary of it, and knowing the boundary is most of what makes it safe to use on documents that actually matter.

Questions people ask

Should I ever let AI write my first draft?
Only when the stakes are near zero and the goal is momentum rather than quality — a placeholder outline you intend to gut. For anything you plan to publish or send, drafting from your own notes is faster once you count the time spent auditing a generated draft for confident-sounding errors.
What should I never ask an AI editing tool to change?
Numbers you haven't verified elsewhere, quotes, and anything legally load-bearing. Ask it to compress, reorder or flatten tone freely; ask it to touch a fact-bearing sentence only after you've confirmed the fact yourself.
How do I stop an AI edit from changing what I meant?
Diff the edited version against your draft line by line before you accept it, reading only for meaning drift, not phrasing. A model asked to "tighten" a sentence will sometimes drop a qualifier or flip an implied causation, and it never flags the change.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.

This article names specific products. How we handle recommendations.