AI transcription tools compared on difficult audio
Five transcription tools on accented speech, crosstalk and bad microphones. Accuracy claims collapse the moment the audio is real.

Part of AI writing in real work, not in demos
Every transcription tool we've ever tried quotes an accuracy figure on its pricing page, and every one of those figures was measured on audio that resembles nothing we actually record: clean single microphone, one speaker, no background noise, standard accent, read from a script or spoken slowly into a quiet room. That's the audio a benchmark needs to be reproducible, but it tells you almost nothing about the transcript you'll get from an interview recorded on a laptop with the fan running, a client call where two people talk over each other in the first minute, or a source with an accent the model was never tuned on. We tested five tools against three samples built to be the audio people actually have, not the audio vendors benchmark against, and the rankings from the clean-audio demos didn't survive contact with any of them.
The three samples, and why each one breaks something different
We recorded three ten-minute samples and ran all five tools against each one, then checked the output against a manual transcript we made by hand and cross-checked between two people.
The first was a solo interview recorded on a laptop's built-in microphone in a room with a running space heater fan roughly a metre from the speaker — the single most common bad-audio scenario for anyone who records without thinking about it in advance. The second was a two-person call with genuine overlap: interruptions, a "sorry, go ahead," and several stretches where both people were talking at once for two or three seconds. The third was a single speaker with a strong regional accent reading the same fifteen sentences every other sample used, so accent was the only variable that changed.
Word error rate roughly tripled across the board between the clean control recording and the three difficult samples, and the tool that came out ahead on the clean control was not the tool that came out ahead on any of the three difficult ones. A benchmark run on easy audio doesn't just understate the error rate you'll see; it doesn't reliably predict which tool makes fewer of the errors that matter to you.
That gap between demo audio and real audio is the same gap that shows up whenever a tool trades a narrow, bounded job for an open one. reach, a CV-to-website generator we've tested separately, works because it refuses to be general — it fills specific, vetted copy slots on a page assembled from an existing résumé rather than writing a page from a blank prompt, generates the result in about twenty seconds, and gets someone live at a free subdomain in under two minutes. The honest cost of that narrowness is that reach produces exactly one page, with no sub-pages and no CMS — a non-issue for a one-page CV site, a disqualifying limit for anything meant to grow into a multi-page site. Transcription has the same shape of tradeoff running underneath it: the models that hold up best on accented or overlapping speech tend to be the ones tuned narrowly for that problem, not the general-purpose ones optimised for the clean demo case, and the same discipline that keeps reach's output reliable is what keeps a transcription model's output usable on audio nobody demos with.
Crosstalk is where every tool loses the most at once
Speaker diarisation — assigning each stretch of speech to the right person — was the single weakest capability across the board, and it degraded worst exactly where accuracy also degraded worst: the overlapping section of the two-person sample. The failure took two different shapes, and both are worse for a usable transcript than an ordinary misheard word, because a misheard word is locally wrong and a diarisation error is wrong about who said something.
The more common failure was merging: two overlapping utterances collapsed into a single turn attributed to whichever speaker's voice was picked up more clearly by the mic, with the other person's words either dropped or folded silently into the same line. The rarer but more confusing failure was the opposite — one person's single continuous sentence, spoken without a real pause, got split across a false speaker change mid-sentence, so the transcript reads as two people finishing each other's thoughts when one person just kept talking. Every tool we tested did the first at least once across the sample; two of them did the second as well. None of the five handled the "sorry, go ahead" interruption pattern — a genuinely common conversational move — without at least one misattribution somewhere in the ten minutes.
Names, jargon and whether a custom dictionary actually helps
This is where the tools separated most clearly, because it's the one problem with a real fix: letting the user tell the tool in advance what words to expect. Three of the five let you upload or type a custom vocabulary list before transcribing — names, product terms, acronyms — and on our jargon-heavy sample that list measurably reduced the number of misspelled proper nouns each time we used it. It did not fix accent-driven mispronunciation of ordinary words; a custom dictionary corrects for words the model has never heard, not for words it mishears because of how they were said. The two tools without a custom vocabulary feature transcribed the same recurring product name a different, wrong way almost every time it appeared, which is a worse outcome for a working transcript than a scattering of unrelated small errors, because it breaks a find-and-replace fix — you end up hunting six variant misspellings instead of correcting one.
Time to a usable transcript is not the same as processing time
Every tool's pricing page leads with processing speed, and processing speed was the least useful number we measured, because the gap between the fastest and slowest tool's processing time was under two minutes on a ten-minute file, while the gap in correction time afterward — the actual work of getting from raw output to something you'd hand to someone else — ran to twenty minutes or more between tools on the same file. The tool with the weakest diarisation on the crosstalk sample was also, unsurprisingly, the one that took longest to correct, because reassigning mislabelled turns by re-listening is slower than fixing individual misheard words in place. Total time to a usable transcript, not raw processing time, is the number worth planning around, and it's the same lesson we kept relearning across AI writing in real work more broadly: the model's own speed is rarely the bottleneck once you count what a human still has to do to the output afterward.
Where the audio goes, and what the terms actually say
Four of the five tools process audio in the cloud; the fifth is an open-weight model that can be run entirely on a local machine, which removes the upload step and any question of where the recording lives afterward, at the cost of needing decent hardware and being noticeably slower on longer files without a capable GPU. For the four cloud tools, retention and training terms were, as we found when we looked at AI meeting notes tools, the hardest thing to get a straight answer on — the reassuring language on a vendor's general privacy page is frequently scoped to a business or enterprise tier rather than the individual plan most people reading a comparison like this one are actually on. If a recording contains anything confidential, checking the terms for your specific plan tier before you upload it is not optional, and it's a different document from the one the marketing page links to.
What the tools cost, set against what they leave you to fix
All five bill by audio duration rather than by seat, which is the sane model for a tool most people use occasionally rather than daily, and all five offer a free tier limited enough to be a trial rather than a real workflow. Beyond that the billing structures diverge — some meter a monthly pool of minutes with overage charges past it, others sell blocks of hours outright — in ways specific enough to each vendor's current pricing page that quoting exact figures here would be stale within a quarter; check the vendor's own page for the number that applies to your volume today. Prices checked August 2026. What we can say from six weeks of running real files through all five: the sticker price per hour of audio was a worse predictor of total cost than the correction time each tool left behind. A cheaper tool that hands you forty minutes of diarisation cleanup on a difficult recording costs more than a pricier one that hands you ten.
Who each one actually suits
If most of your audio is a single clear speaker or a well-behaved two-person interview without much overlap, the differences between these tools shrink close to the point of not mattering, and the deciding factor becomes whether you need a custom vocabulary list for names and jargon — which we'd treat as close to mandatory if you're transcribing anything with recurring product names or people's names that a general model has never seen.
If your work regularly involves overlapping speech — panel discussions, group calls, anything with more than two people who interrupt each other — diarisation quality matters more than raw word accuracy, because a merged or misattributed turn corrupts the transcript's usefulness in a way that a handful of wrong words doesn't. Budget the correction time accordingly rather than trusting the processing-speed number on the pricing page.
If confidentiality is the binding constraint — client interviews, anything under an NDA, source material you don't want anywhere near a vendor's training pipeline — the local, open-weight option is the only one of the five that removes the question entirely rather than asking you to trust a plan-tier-specific promise. The tradeoff is genuine: slower on long files without decent hardware, and no hosted vocabulary management to lean on, so you're doing more of the setup yourself in exchange for the recording never leaving your machine. For anyone whose primary difficult audio is a strong accent rather than crosstalk or confidentiality, we'd point instead at whichever tool's custom vocabulary list is most flexible about phonetic spellings, since accent errors we saw were concentrated in specific recurring words rather than spread evenly across the transcript, and a good vocabulary list is the one lever that reliably brings that specific error type down. We go through the editing side of this same pattern — where a model's fluency helps and where it papers over an error you'd otherwise catch — in the writing tools we tested against real documents, and the overlap with transcription is closer than either category likes to admit: fluent, confident, and wrong in ways that don't announce themselves is the exact failure mode both share.
Questions people ask
- Which AI transcription tool is most accurate?
- On the kind of clean, single-speaker audio vendors demo with, the differences between leading tools are small enough not to matter. On real audio — accents, crosstalk, a laptop mic in a room with a fan — the ranking changes completely, and the tool that wins depends on which of those problems your recordings actually have.
- Do transcription tools handle regional accents well?
- Unevenly, and the gap between the best and worst tool we tested on a single strong regional accent was larger than the gap between any two tools on clean audio. A custom vocabulary list narrows this for jargon and names but does not fix pronunciation-level accent errors.
- Can AI transcription separate two people talking over each other?
- Better than it used to, but overlapping speech is still where every tool we tested lost the most accuracy at once, and the failure mode — merging two speakers' words into one turn, or splitting one continuous sentence across a false speaker change — is worse for the transcript's usefulness than a few wrong words scattered through clean passages.
- Is it faster to run transcription locally or use a cloud tool?
- Locally-run open models remove the upload and processing wait entirely, which matters most for long files, but the correction pass afterward takes the same amount of a person's time either way, and that correction pass is usually the larger share of the total time to a usable transcript.