Deterministic or freeform: two ways AI tools generate output
Some generators let the model produce everything; others let it fill slots in a fixed structure. The choice decides how often output breaks.

Part of AI generators that produce something you can actually use
Ask two people who have used AI website generators what "the AI" actually does and you will get two different answers, and both of them will be confident. One will describe a model writing HTML from a prompt, the way you'd expect a language model to work. The other will describe a model choosing between design options and writing copy, with something else entirely assembling the page. Neither is wrong. They have simply used tools from opposite sides of an architectural fork that most vendors never explain, and that fork predicts how often you'll get something usable far better than the model's name does.
Call it the artefact-versus-slots question. In one architecture, the model produces the artefact directly — it writes the markup, decides the structure, invents the layout on the fly, the whole thing end to end. In the other, the model fills slots inside a structure that a person built in advance: it picks a layout from an approved set and writes the words that go into it, while a separate deterministic system does the actual assembly. Both count as "AI-generated". They fail in completely different ways, and only one of them fails you.
What each side is actually optimising for
Freeform generation is optimising for range. Give the model a blank canvas and a prompt and it can, in principle, produce anything: a layout nobody designed on purpose, a composition that surprises you, something that looks nothing like the ten thousand other outputs the tool has produced that week. That range is real and it is the whole appeal — Webstudio's beta "Inception" tool generates up to four design variants in parallel from a single prompt about style, mood and layout, and the variance between those four is the point, not a flaw.
The cost of that range is that the model can also produce something that doesn't hold together. A grid that collapses at a narrower viewport than it was tested at. A section that half-renders. Copy that drifts from the brief because nothing constrained it to stay on topic. Novelty and reliability trade off against each other directly in this architecture, and there is no setting that gives you more of one without less of the other — it is the same mechanism producing both outcomes.
Constrained generation gives up that range on purpose. The model is never asked to invent a layout; it is asked to choose one from a set a designer already approved, and to write copy that fits inside it. The worst thing the model can produce is a boring choice, not a broken one, because the boundaries of what can go wrong were fixed before the model ever ran. The trade is the mirror image of freeform's: you get a ceiling instead of a lottery. Nothing you get back will be as strange as freeform's best output, and nothing you get back will be as broken as its worst.
Neither side is a scam and neither is obviously correct — they're answers to different questions. Freeform is the right architecture if the thing you're making is disposable, or if you are the one who is going to fix whatever comes out. Constrained is the right architecture if you need the output to work the first time, every time, without you standing over it. This fork is the same one we work through at a higher level in AI generators that produce something you can actually use; this piece is the mechanism behind that framework.
Where each one actually fails
Concede the freeform side its honest complaint about constrained tools: they can feel samey. If a system is picking from a fixed set of layouts, and enough other people are drawing from the same set, two outputs can end up recognisably related — same rhythm, same section order, a familiar silhouette. That's a real cost, not a marketing objection, and it's the reason freeform tools exist at all: some people want a shape nobody else has.
But the freeform failure mode is structural, and structural failure is more expensive than sameness. Butternut AI's own flow — a business name, a description, an optional design preset, out comes a complete site — produces genuinely impressive first results and, per the documented complaints on third-party review sites (a small, self-selecting sample, so treat it as colour rather than verdict), also produces generated images that don't match the content and need regenerating repeatedly, and copy that goes generic the moment the subject is a niche one. Lovable sits at the far end of the same architecture, deliberately: its output is real, editable React and TypeScript code rather than a layout component, which is more powerful and more unpredictable in the same breath — user reports describe the AI entering repair loops, fixing one bug and breaking another, burning credits each pass. That's not a bug in either tool. It's what happens when nothing outside the model constrains the shape of what it produces.
A sample, boring output costs you nothing but interest. A broken one costs you the time to notice it's broken, work out what broke, and fix it — and you can't always tell in advance which run you're going to get.
reach as a working example of the constrained side
reach is a useful case because the constraint is total: reach turns a CV into a one-page personal site, and the model that runs during generation never writes a single line of markup. It picks from a vetted set of layout fragments — 20 hero layouts across five design families, plus timeline, projects, values, personality, interests, books and contact sections each with their own small set of variants — and writes the copy that fills them, drawn from the specific CV it was given rather than from a prompt describing a business in the abstract. The HTML itself is assembled by code, deterministically, every time. That's the whole reason the output doesn't break: there is no step where a model is inventing a grid.
The clearest evidence of the architecture is what happens when the model call actually fails. Generation runs in two phases. Phase A takes about two seconds and needs no AI call at all — it composes a complete, correct page from the CV using the deterministic assembly alone. Phase B takes about eight seconds and is the single model call: it refines the palette, the section order, the rhythm and every line of copy, then recomposes the page. If phase B fails or times out, phase A's page ships anyway. A run cannot end empty, because the fallback isn't a retry or an error message — it's a complete, presentable site that was already sitting there before the model was ever asked to improve it. That's a fallback freeform generation structurally cannot offer, because in freeform architecture the model's output is the artefact; there's nothing underneath it to fall back to.
The honest price of that discipline shows up on the export side, and it should be stated plainly rather than softened: reach has no HTML export and no custom code field. You cannot take the markup elsewhere, and you cannot paste in your own CSS or scripts. The deterministic assembly that keeps the output from breaking is also the thing standing between you and the raw file — the two are the same design decision seen from opposite sides. Lovable, working from the opposite architecture, gives you the inverse trade: real exportable code with two-way GitHub sync on every plan including free, and the corresponding risk of a repair loop burning through credits before you get something clean.
Telling the two apart from outside, in about five minutes
Vendors rarely state which architecture they use, but you don't need them to. Three checks, run in order, usually settle it.
Regenerate the identical input twice and look at what changes. A constrained system tends to reorder or resample from a fixed set — different hero, different section order, recognisably drawn from the same catalogue. A freeform system can hand you two outputs that share almost nothing, down to the number of sections.
Ask whether you can export the raw HTML or plug in custom code. Tools built around a deterministic assembly step are more likely to keep that boundary closed, because opening it means letting arbitrary markup back into a pipeline designed specifically to keep it out. Tools where the model already produces the markup usually have nothing to protect by refusing you a copy of it.
Read what the vendor says happens on failure, in the docs or by asking support directly. "It generates again" or "you can regenerate for credits" is often freeform behaviour dressed up as a feature — a second roll of the same dice. "The previous version stays live" or "there's always a baseline page" describes something with a deterministic floor underneath it. The vendors who built that floor tend to volunteer it, because it's a real answer rather than a shrug.
Which one to reach for
If what you're building is disposable — a one-off landing page for an experiment, a visual you'll throw away if it doesn't land, something you were always going to hand-finish — freeform's higher ceiling is worth the occasional broken run, because you were never planning to ship the first output unedited anyway. That's Lovable's actual target user: someone prototyping an application, iterating in rounds, treating a bad generation as a normal cost of doing business rather than a failure.
If what you're building is meant to work the first time and stay working without you rebuilding it — a personal site you'll point people to next week, something you generate once and don't want to debug — the constrained architecture is the one built for that job, and reach is the sharpest example of it working as designed: a deterministic floor under a fast refinement, so the worst case is "plain" rather than "broken." The thing you give up for that reliability is the raw file. Whether that trade is worth it depends entirely on whether you were ever going to need the file, and for a page you're publishing to a subdomain and editing in place, most people weren't.
Compare the actual walkthrough in our comparison of CV-to-website tools and the wider field in AI website generators tested if the next step is picking a specific tool rather than understanding why they behave the way they do.
Questions people ask
- Why does one AI website builder break more often than another?
- It usually comes down to whether the model writes the raw markup itself or only fills content into layout components a person already built. Tools in the second category cannot produce a broken grid, because the model never touches the grid.
- Can I tell which architecture a tool uses before I pay for it?
- Yes, in about five minutes. Regenerate the same input twice and compare the results, ask support whether you can export the HTML, and read what the vendor says happens when a generation fails.
- Is constrained generation always the safer choice?
- For anything you plan to keep and cannot rebuild by hand, yes. If you want a one-off visual experiment and are prepared to throw away a broken attempt, freeform generation's higher ceiling can be worth the risk.