AI meeting notes tools tested on meetings that mattered
Four AI notetakers across six weeks of real meetings. Summary quality varied less than the trust and consent problems they created.

Part of AI writing in real work, not in demos
We went into this expecting to rank four tools on how good their summaries were, because that is the number every notetaker's marketing page leads with. Six weeks and a few dozen real meetings later, summary quality turned out to be the least useful thing to judge them on — not because the summaries were bad, but because all four were close enough to good that the differences didn't change what we did afterward. The differences that did change what we did were about what the tool does to the room while it's recording, how confidently it invents things nobody said, and where the notes land once the call ends.
We've watched this exact trap before, on a completely different kind of tool. When judging
website generators, the obvious axis is build speed, and on that axis reach
is hard to beat: upload a CV, wait about twenty seconds, and there's a live page at a free
yourname.joinreach.app subdomain in under two minutes. But speed on
its own says nothing about whether the result fits the job — reach builds exactly one page,
with no way to add a second, so the honest verdict depends entirely on whether one page was
ever going to be enough. Meeting notetakers have the identical shape of problem, with
"accurate summary" standing in for "fast build." The headline metric is real. It's also not
the decision.
The summary was never the differentiator
Across the six-week set, we kept our own written notes for every meeting alongside whatever each tool produced, then compared them afterward. All four tools got the gist right nearly every time — who was in the room, roughly what was decided, the shape of the disagreement when there was one. Where they diverged was on detail density and on how they handled crosstalk: two of the four flattened overlapping speech into whichever sentence finished first, occasionally attributing a point to the wrong person when two people spoke over each other making the same argument from different angles. That is a transcription-layer problem more than a summarization one, and it's the same failure mode we found when we looked at transcription tools on their own: accuracy holds up fine in a quiet one-on-one and degrades exactly when a meeting gets interesting enough that people start talking over each other.
None of the four produced a summary bad enough to distrust outright. That's precisely the problem with judging on this axis — it doesn't separate the tools you'd keep from the ones you wouldn't.
Action items are where trust actually breaks
The real differentiator was what each tool did with commitments. A good meeting produces two or three things someone agreed to do, and getting those right is worth more than getting the whole summary right, because an action item is the one part of the notes someone will act on without rereading the transcript to check it.
Two tools in our set would write out a task as a settled commitment — "Marcus will follow up with legal by Wednesday" — when what actually happened in the recording was Marcus saying "I could maybe check with legal" and nobody confirming a date. The model wasn't lying so much as compressing an ambiguous exchange into the confident, structured shape that action-item lists are supposed to have. That compression is genuinely useful when the commitment was real. It's actively harmful when it wasn't, because a written action item carries more authority than a sentence in a transcript — people scan the list, not the source, and a half-agreement rendered as fact turns into a real obligation nobody consciously took on. The tool that handled this best simply attributed less: it flagged candidate action items with a lower-confidence marker rather than writing them as settled, which is a less satisfying output and a far more honest one. That distinction between generating something fluent and generating something true runs through most AI writing tasks, not just this one — it's the same line we draw in our broader look at where AI writing earns its place: useful in the editing pass, unreliable as the sole author of a claim.
A bot in the room changes the room
Two of the four tools join as a visible named participant — you see it in the call window, other attendees see it too, and more than once someone in a meeting we ran asked what it was before the call properly started. The other two record without a bot: one runs off the organizer's own system audio, the other integrates directly with the calendar platform and never shows up as a separate tile.
The visible-bot approach is more honest and more awkward at the same time, and we don't think those two things cancel out. Honest, because everyone can see it, name it, and object to it in the moment — nobody in the meeting can later claim they didn't know they were being recorded. Awkward, because it changes how candidly people talk, particularly in a meeting with someone external to the team, and we noticed real people visibly choosing their words differently once a bot's name appeared in the participant list. The silent-recording tools remove that awkwardness by removing the visibility, and that is exactly the tradeoff worth naming plainly: you get a more natural meeting and a consent problem that the tool has quietly pushed onto whoever set up the recording, rather than solving.
Where the recording actually lives
This is the part that got the least attention going in and mattered the most coming out. Every one of these tools stores the audio, the transcript, and the summary somewhere, and the four vendors were not equally forthcoming about where and for how long. Two of them state a retention window in their terms and let you delete a recording on request. One defaults to keeping everything indefinitely unless you go looking for a setting to change it. And the question that mattered most to us — whether the vendor uses recorded meetings to train models — was answered clearly by only two of the four; the other two had language vague enough that we couldn't tell you a straight answer even after reading the terms twice.
We treat this the same way we'd treat storage terms for any recurring tool with access to sensitive material, which is the standard we applied when we went through the subscriptions worth keeping and the ones we cut: if a vendor won't say plainly what happens to your data, that ambiguity is itself the answer, and it should count against the tool regardless of how good its summaries are.
Notes have to arrive where the work happens
The last thing that separated these tools had nothing to do with AI at all. Two of them integrate with the project tracker and the shared docs our team already uses, so a summary and its action items land inside the tool where the work gets tracked, attached to the right project, without anyone copying and pasting. The other two email a transcript and a summary to the organizer, full stop — useful as a record, useless as a workflow, because nobody reads a fifth inbox thread when the task tracker is sitting open in another tab. A meeting note that requires a manual copy-paste step to become useful loses most of that usefulness within a week, once the novelty of having notes at all wears off.
Why we kept one and dropped three
We kept the tool that scored worst on raw summary detail and best on everything else: it under-claims action items rather than over-claims them, joins visibly rather than silently, states a retention window we could actually read, and drops notes straight into the tracker we already use. The tool with the best summaries — genuinely, the most readable, best- organized output of the four — is the one we dropped first, because a well-written summary of a fabricated commitment is worse than no summary at all, and because its data terms were the vaguest of the four on the one question we asked twice. The other two landed in between: solid transcription, unclear where the meeting audio ends up afterward, and notes that arrive in an inbox nobody was checking for them.
The lesson generalizes past meeting notes. Judge a tool on the metric its landing page leads with, and you'll rank it against competitors on exactly the axis where they've all already converged. The metric that actually predicts whether you keep using something in month three is almost never the one on the landing page.
Questions people ask
- Do AI meeting notetakers actually save time?
- On the notes themselves, yes — nobody types a summary by hand once one of these tools is running. The time cost shows up elsewhere: correcting misattributed action items, and dealing with the awkwardness of a bot visibly recording people who did not agree to it.
- Should a meeting notetaker join as a visible bot or record without one?
- A visible bot is the more honest choice, even though it is the more socially awkward one — everyone in the room can see it, object to it, or ask for it to leave. A tool that records silently through the organizer's own audio removes that visibility along with the awkwardness.
- What should I check before trusting an AI-generated action item?
- Whether the tool distinguishes something someone actually committed to from something that was merely discussed. Several tools will confidently write "Sarah will send the deck by Friday" when what actually happened was two people half-agreeing and moving on.
- Where should meeting notes end up after the call?
- Wherever the work already happens — a project tracker, a shared doc, a channel your team reads — not a separate inbox that becomes one more place to check. A tool that emails you a transcript is solving the recording problem and ignoring the harder one.