Cited

How we test a tool before writing about it

Our testing method: a fortnight of real use, a fixed brief per category, and the questions we answer before a verdict goes out.

A vintage pocket watch rests on an open antique book surrounded by dried yellow roses.
Photo: Ylanite Koppens / Pexels

You can tell within a paragraph when a review was written from a trial account opened that morning. The praise is generic, the screenshots show the empty-state onboarding flow, and the "verdict" is really just a tour of the pricing page with adjectives attached. We know this because most of what we used to publish read exactly like that, and it is why we rebuilt the process from the ground up rather than fixing it in place.

The honest complaint against what follows is that it is slow. A two-week minimum per tool means we cover fewer things than a site that turns around a "review" the same week a product launches. We accept that trade on purpose. A tool that looks identical to five competitors on day one usually stops looking identical by day ten, and day ten is the only day that matters to someone who is about to pay for a year of it.

The fortnight rule

Every tool we write about gets used for real work for a minimum of two weeks before a word gets published. Not explored — used. That distinction is the whole method.

A one-day trial catches the interface: is the onboarding clear, does the first export work, does the free tier feel stingy. What it cannot catch is the stuff that actually determines whether the tool earns its subscription: does an automation someone built in week one still run correctly in week three after an upstream field changed. Does the AI output degrade or drift once you are past the demo prompt and into your actual, messier inputs. Does the tool's idea of "unlimited" survive a week where you genuinely push volume through it. Does support answer a real ticket, or only the kind of question their own docs already cover.

None of that shows up on day one, because day one is the day every vendor has optimized for. A fortnight is roughly the point where the tool stops performing for a first impression and starts behaving the way it will behave for the rest of your relationship with it — which is also, not coincidentally, close to the point where a lot of tools we have tested in the past quietly develop the friction their marketing never mentioned.

The same brief, every time, within a category

Comparing five automation platforms is meaningless if each one is tested against a different task someone happened to make up on the spot. So within a category, every tool runs the identical brief. If we are testing automation that survives a month, every candidate builds the same workflow against the same trigger conditions and the same edge cases, not whatever demo the vendor's own template gallery happens to make easy. If we are testing generators, every candidate gets the same input material and the same set of outputs we ask for, described in the piece on AI generators that produce something usable.

This is less interesting to read about than it is to actually do, because it means we often can't use a tool's best showcase feature if a competitor doesn't have an equivalent, and it means some tools look worse on our brief than they would on a brief they designed themselves. That is deliberate. A comparison where each product gets to pick its own test is not a comparison, it is five separate advertisements laid side by side.

The failure questions

Every review answers three questions that have nothing to do with the feature list, because the feature list is the thing every vendor already tells you.

What happens on the worst day. Not the demo day — the day the API you depend on is down, the day you fat-finger a bulk edit, the day your file is larger or messier than anything in the tutorial. Does the tool degrade gracefully, throw a comprehensible error, or silently produce wrong output that looks right. We have killed recommendations over this alone: a tool that is pleasant in the sunny-day case and dangerous in the rainy one is worse than a plainer tool that just tells you it failed.

What happens at the price ceiling. Every tier has a wall somewhere — a usage cap, a seat limit, a feature gate — and the number on the pricing page rarely tells you what it feels like to hit it mid-task. We test at the edge of whatever tier we're reviewing on purpose, not just below it, because the moment you discover a ceiling exists is usually the moment you're already depending on the thing behind it.

What happens on the way out. Can you get your data, your content or your work out in a usable form, or does leaving mean starting over somewhere else. We have found this to be the single most under-reported fact in the entire category. A tool with a mediocre editor and a clean export is a safer bet, long term, than a beautiful editor that locks your work inside its own format.

We also note where an ecosystem has security exposure that the product page won't mention, when a documented figure exists to cite. WordPress's own extension ecosystem is a genuine example: Patchstack's State of WordPress Security in 2026 recorded 11,334 new vulnerabilities across the ecosystem in 2025, a 42% rise year on year, with 91% of them in plugins rather than in WordPress core. That is not a reason to avoid WordPress — it is a reason to write "the risk sits in what you install, not in the software itself" instead of leaving the question unaddressed.

What we don't measure, on purpose

We do not run feature-count tables — the "47 features versus 52 features" format that looks rigorous and measures nothing. Feature counts reward vendors who ship a toggle for everything and say nothing about whether any individual feature works well, is discoverable, or is worth the complexity it adds to the interface. Two tools can share a feature list and be completely different products to actually use.

We do not publish star ratings. A single number implies the tool is equally good for every kind of buyer, and the honest answer is almost always "good for this kind of work, wrong for that kind" — which a rating collapses into noise. Where a genre-appropriate review score exists elsewhere, we say so and name the sample size, because a rating from fifteen reviews is a different kind of evidence than one from fifteen thousand, and treating them the same is its own small dishonesty.

We do not chase every new release the week it ships. An AI feature announced on launch day is often materially different from the same feature three months later, once the vendor has had time to fix the parts that broke in production. We would rather be a month behind and accurate than first and wrong.

Disclosure

We do not accept vendor-supplied test accounts, comped subscriptions, or review units of any kind. Every account we use to test a paid tool is one we signed up and paid for ourselves, the same way any other customer would. We run no sponsored reviews and no affiliate-led shortlists — nothing on this site is included, excluded or ranked because of a commercial relationship with the company that makes it. If that ever changes for a specific piece, it will say so at the top, not in a footer.

This also means our coverage has gaps. There are tools we simply have not paid for yet, categories we have not built a brief for, and products that launched last week that we will not have an opinion on for another two. We would rather publish less and mean all of it than fill the gaps with a rewritten press release, which is most of what happens in this category and most of why the ai tool graveyard exists in the first place — a fair number of the tools in it got glowing first-week coverage from outlets that never went back to check.

Reproducing this yourself

None of this requires access we have and you don't. Pick the category, write down the same brief for every candidate before you touch the first one, and hold yourself to using each tool for actual work — not a tour — for at least two weeks before forming an opinion. Ask what happens when it fails, what happens when you hit the ceiling of the tier you're paying for, and what happens the day you try to leave. Those three questions will tell you more than any comparison table, ours included, and they cost nothing but the patience to wait a fortnight before writing the sentence you already wanted to write on day one.

Questions people ask

How long do you actually use a tool before reviewing it?
A minimum of two weeks of real, task-based use inside a fixed brief for that category, not a single afternoon with a trial account.
Do you accept free accounts or review units from vendors?
No. Every account we test with is one we signed up and paid for like any other customer, and we do not accept vendor-supplied logins.
Why don't your reviews include a feature checklist or star rating?
Because a feature list tells you what a tool claims to do, not what happens when you actually need it to do that thing on a bad day.
Can I run the same test on a tool myself?
Yes — the brief, the failure questions and the categories are described below specifically so you can reproduce them and check our conclusions.

Cited — We use a tool for a fortnight before we write a word about it, and we say where every number came from.