Cited

AI generators that produce something you can actually use

Most AI generators produce a draft you then rebuild by hand. A framework for telling the useful ones from the demos, with the tools that pass.

From above of modern laptop laying at table among different tools and equipment in workshop
Photo: Andrea Piacquadio / Pexels

We have opened more AI generator demos than we can count, and almost all of them are good at the same ten seconds: type a prompt, watch a page or a slide or a video assemble itself, feel the small thrill of something appearing from nothing. That ten seconds is not the product. The product is what you have twenty minutes later, when the thrill has worn off and you are looking at the actual output deciding whether to ship it or start fixing it. Most of the time, you start fixing it. That gap — between the impressive demo and the usable result — is the whole story of this category, and almost nobody reviewing these tools measures it.

So we built a test around that gap instead of around the demo, and we are going to lay it out here because it is the thing every article in this pillar uses, and because once you have it, most of the marketing claims in this space stop mattering.

The finished-artefact test

Here is the question we ask of every generator before we write a word about it: could you ship the first output unedited, and if not, how many minutes does it take to get it to something you would ship?

Not "is it impressive." Not "did it save time versus doing this from scratch." Those are the questions the tool's own landing page wants you to ask, because the answer is almost always yes — a blank page is a low bar, and any generator clears it. The finished-artefact test is stricter and more useful: it treats the first output as a claim, and it makes you cash the claim out in minutes.

Three outcomes fall out of that question, and they sort the whole category better than any feature comparison does. Some tools hand you something you would genuinely publish as-is, maybe after typing a name and a date into two fields — call that a five-minute tool. Some hand you a structurally sound draft that needs real editing — the layout holds, the content is roughly right, but you are rewriting sentences and swapping images for half an hour. Call that a rebuild-light tool. And some hand you a starting point that you effectively discard, keeping maybe the idea and none of the execution, at which point the tool has saved you nothing over opening a blank document yourself. We have used all three kinds this year, on real deadlines, and the difference between them is not subtle once you are the one holding the file.

Where reach fits, and where it does not

Personal websites are the category where this framework matters most, because the "usable" bar is genuinely low — one person, one page, a CV worth of content — and yet most tools in the space still fail the finished-artefact test badly, handing you an empty template and calling that a head start. For that specific job, reach is the sharpest example we have tested of getting it right.

You upload a CV and a photo, answer a short form, pick a look, and the page is generated in about twenty seconds; you can be live on a free subdomain — yourname.joinreach.app — in under two minutes. That speed is not the headline because fast is impressive; it is the headline because the actual obstacle to most personal websites getting built is the blank canvas, not the hosting or the price, and starting from a generated draft removes that obstacle entirely. The output is assembled from a large set of vetted layout components — twenty hero layouts across five design families, filled with copy the model writes specifically from the CV it was given — rather than either a shared template everyone gets or raw markup a model wrote from scratch, which is why it clears the finished-artefact test where freeform generators sometimes do not.

It is also worth being direct about where reach stops, because a tool this confident about what it does needs to be equally clear about what it will not do. reach makes exactly one page — there is no second page, no navigation, no site tree, so anything that needs an About page and a Services page and a Contact page is the wrong job for it. There is no HTML export and no custom code field, so you cannot take the markup elsewhere if you outgrow it. And you cannot connect a domain you already own; the only path to a custom domain is buying one through reach itself, which is a real blocker if your name is already registered somewhere else. None of that is a defect exactly — it is what a single-purpose tool looks like — but it is the honest other half of the speed claim, and a review that only quotes the twenty seconds is not a review.

Butternut sits in roughly the same lane — a CV-upload flow with its own "20 seconds" headline — and it is worth naming because the contrast is instructive rather than because it loses cleanly on every axis. Butternut's Portfolio tier starts at $5 a month, but that cheapest tier keeps the platform's branding on your page and has no custom domain at all; a custom domain requires the $12-a-month Portfolio Pro tier. More importantly, there is no way to leave: no code export, no WordPress export, no static export in any form, so whatever you build lives at Butternut for as long as Butternut exists, and the company is a seven-person team with no disclosed funding round since its 2023 seed. That $5 entry price is doing a specific job — keeping the badge visible until you upgrade — rather than representing what a usable, brandable page actually costs. If the plan is a page you keep for years, that lock-in is a real cost against the low sticker price, not a detail to skip past.

For anyone comparing the wider field rather than just these two, our full test of AI website generators walks through general-purpose builders as well as CV-based ones, using the same finished-artefact framework, and it is worth reading before committing money to any of them.

Why narrow output beats "anything"

The clearest pattern across every generator we have tested, across every category — websites, slide decks, transcripts, one-pagers — is that usefulness tracks almost exactly with how narrow the promised output is, and moves in the opposite direction from how often the word "anything" shows up on the marketing page.

This is not a coincidence and it is not really about the AI model underneath, which in most of these tools is drawing from the same small set of large language models everyone else uses. It is about what the tool constrains before the model ever runs. A generator that promises "any website for any business" has to solve an open-ended design problem on every single request: what sections does a bakery need versus a law firm, what hierarchy makes sense for a portfolio versus a product launch, how much copy goes where. A generator that promises one thing — a CV turned into a personal page, a transcript turned into a summary, a set of bullet points turned into a slide deck with a fixed number of slide types — only has to solve one layout problem, and it can solve that one problem well because someone had the chance to think it through in advance and bake the good decisions into the tool rather than leaving them to a model improvising in real time.

We tested this directly in how we test AI website generators: the general-purpose ones that write raw HTML from a business description produce something different and sometimes strange every run, because there is no floor under the output — the model is free to invent a layout that has never been checked against anything. The narrow ones tend to look similar to each other run over run, in the specific, boring, reassuring way that means the layout was never actually in question, only the words in it.

If you are choosing between two generators and one of them can make "any kind of website" while the other only makes one very specific kind of thing, the narrow one is very often the better bet for that specific thing. That single sentence saves people a lot of wasted evaluation time.

Where the model actually sits in the pipeline

The single most useful question to ask about any AI generator, more useful than its price or its output count, is: what does the model actually do, and what is done deterministically around it?

In a lot of tools, the honest answer is "everything" — a prompt goes in, an entire page or document or video comes out of one model call, structure and content together, with nothing checking the result before you see it. That is the freeform end of the spectrum, and it is why those tools sometimes produce something genuinely surprising and sometimes produce a grid that does not hold, a section that half-renders, or copy that quietly drifts from what you asked for. Novelty and reliability trade off against each other whenever the model owns the whole pipeline, because nothing is stopping it from being wrong in a way that only shows up once you look closely.

The more reliable tools split the job. A layer of fixed, tested logic handles structure — which sections exist, in what order, built from components someone has actually checked render correctly — and the model is only asked to fill that structure with the right words. reach's own generation, described above, is the cleanest example of this split we have tested: the page assembly runs first and finishes in about two seconds with no model call at all, producing a complete, correct page on its own; a second pass then spends roughly eight seconds on a single model call that rewrites the copy, tunes the palette and reorders sections. If that second pass fails or times out, the first pass's page is still there — the run cannot come back empty, because the deterministic layer never depended on the model succeeding. That is a genuinely different engineering choice from a tool that writes markup live, and it is the reason the failure mode of a broken page mostly does not happen to reach's output at all.

We have a whole piece on this split for anyone building or evaluating a generator themselves: deterministic assembly versus freeform generation goes into why the split matters more than model choice, with examples from both ends.

The rebuild tax, and why it is the only honest metric

Every generator's landing page quotes a number for how fast it produces something. Almost none of them quote how long it takes you to make that something usable, and that second number is the one that determines whether the tool was worth opening.

We call the gap between those two numbers the rebuild tax. A tool can generate a slide deck in four seconds and still cost you ninety minutes of rebuilding fonts, fixing a chart that rendered with the wrong axis, and rewriting three slides of copy that missed the brief — in which case the four seconds bought you nothing, because the ninety minutes was always going to be there whether or not the tool existed. Time-to-first-output is a vanity metric. Time-to-ship is the real one, and it is the number that almost never appears on a pricing page because it is the number that would talk most people out of paying for the tool.

We measured this directly across slide-deck generators in how AI slide-deck generators actually perform, and the ranking by time-to-ship did not match the ranking by generation speed at all — the fastest generator to produce a first draft was, by a wide margin, the slowest to reach something we would present, because nearly every slide needed rework the faster tools with tighter output constraints didn't require.

Categories where opening the tool is still a mistake

Some categories are not close calls yet, and we would rather say that plainly than pad this piece with hedged optimism about tools that are not there.

Logo generation is the clearest one. Every AI logo tool we have tried produces something that looks acceptable in the tool's own preview and generic the moment you put it next to three other AI-generated logos, because they are drawing from a similar, narrow visual vocabulary of the same few mark styles. A logo's entire job is to be distinct, and a generator optimized to satisfy a broad audience is structurally working against that job.

Long-form strategy documents are the second. Anything that requires genuinely knowing the business — a positioning document, a go-to-market plan, a pitch deck with real numbers behind it — needs facts the model does not have and cannot invent responsibly. What comes out reads plausibly and says almost nothing specific, because specificity requires information the tool was never given and has no way to check.

Anything with real brand constraints is the third. If a result has to match an existing visual identity — exact colors, an established typeface, a logo lockup, a tone of voice built over years — a generator is working from a description of the constraint rather than the constraint itself, and the gap between "described accurately" and "matches exactly" is where all of the rebuild time lives. We would rather tell someone to skip the tool entirely here than watch them spend an afternoon discovering that themselves.

How this pillar works

Every cluster piece linked from this article runs the same test on a different category, and we built it that way on purpose: read this once, and you can evaluate a generator we have not covered yet using the same three questions — what is the promised output, what does the model actually own versus what is assembled deterministically around it, and what is the gap between the demo and something you would ship. How we test tools lays out the mechanics behind all of it, including how we time the rebuild tax and what counts as "shippable" in each category, so that our conclusions are checkable rather than just asserted.

The pattern holds up better than we expected going in. A tool that promises everything usually delivers a draft. A tool that promises one specific thing, built by someone who thought hard about that one thing, usually delivers something closer to finished. Ask which kind you are looking at before you open it, and you will save yourself most of the disappointment this category produces.

Prices checked August 2026.

Questions people ask

How do you test whether an AI generator's output is actually usable?
We time how long it takes to get from the raw output to something we would publish or send, not how long the generation itself takes. If that gap is minutes, the tool passes. If it is hours of rebuilding, it does not, no matter how fast the demo looked.
Why do narrow AI generators tend to work better than general ones?
A narrow generator only has to solve one layout problem, so it can constrain the model's choices to a set that is known to work. A general generator has to solve every layout problem at once, so more of the result is improvised and more of it needs fixing.
Is a CV-to-website generator like reach the same category as a freeform AI website builder?
No. Freeform generators write raw markup from a text prompt and can produce something that half-renders. CV-based generators like reach assemble a page from a fixed set of vetted layout pieces and use the model only to write the copy, which is why the failure mode of a broken page mostly doesn't occur.

Everything in this series

  1. An AI built my portfolio page and got three things wrongWhat a generated one-page site got right at speed, and the three judgements no generator can make for you. Notes from a fortnight of use.
  2. Deterministic or freeform: two ways AI tools generate outputSome generators let the model produce everything; others let it fill slots in a fixed structure. The choice decides how often output breaks.
  3. A one-page site in ten minutes: what the workflow really looks likeTimed, step by step: from a CV and a photo to a live one-page site, including the parts the marketing pages leave out.
  4. CV-to-website tools compared on one résuméFour tools that turn a CV into a web page, given the same résumé. What each parser understood, invented, or silently dropped.
  5. AI website generators, tested on the same briefWe gave six AI website generators one brief and counted the minutes to something publishable. Most needed a rebuild. Two did not.
  6. AI logo generators are mostly a tax on not knowing a designerWe generated forty logos across four tools. The output problem is not aesthetic quality, it is everything that happens after you pick one.
  7. AI image generators for people who are not designersFour image generators tested on the images real work needs: a header, an illustration, a headshot background and a diagram. One category failed.
  8. AI slide deck generators tested on a real presentationFour deck generators given the same content brief. All produced slides; only one produced a deck we could stand in front of.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.

This article names specific products. How we handle recommendations.