Here is the short version: record one real meeting, run the same audio through MeetNotes and through whatever meeting app you use today, delete both product names from the outputs, and paste them into ChatGPT, Claude or Gemini with the judge prompt below. Whatever it says is the answer. We will publish the ones we lose, and we will send you 300 free minutes either way.
That is the whole thing. The rest of this post explains why we are doing it that way instead of writing another "we're the most accurate" claim that nobody has any reason to believe.
The problem with every accuracy claim, including ours
Go and search for the best AI meeting transcription app. You will find twenty ranked lists. Nearly every one of them is published on the blog of a company that sells one of the entries — and the entry that wins is, by an astonishing coincidence, the one the publisher sells.
Now look at the accuracy numbers. "99% accurate." "Industry-leading WER." These are real measurements taken on real datasets, and they are also close to meaningless for you, because they were measured on clean studio audio of one American speaker reading prepared text. Your meeting is six people in a room with a ceiling fan, two of whom are talking at once, and at least one of whom switches language halfway through a sentence.
We could publish our own number. It would be just as unfalsifiable as everyone else's, and you would be just as right to ignore it. So instead of asking you to trust a number, we are handing you the test.
The protocol
One. Record a real meeting. Not a scripted demo, not a podcast — a meeting with interruptions, accents and cross-talk. Twenty minutes is enough. Keep the audio file.
Two. Run that same file through MeetNotes and through the other tool. Same audio, same day. One attempt each. If you re-record because the first result was bad, you are testing the meeting, not the tool — and that rule binds us exactly as hard as it binds them.
Three. Strip out every product name, logo and footer from both outputs. Label them A and B. This step is not optional and it is not a formality, which we will come back to in a moment.
Four. Paste both into an AI you already use, with the seven-criterion judge prompt. Ask it to be blunt. Run it in two different models if you want to know whether the gap is real or just one model's taste.
Why the names have to come off
This is the part most "AI-judged comparisons" get wrong, and it is the reason so many of them are worthless.
Large language models have read the internet. The internet contains an enormous quantity of marketing copy, funding announcements and review-site praise about the big incumbent brands, and comparatively little about a small app from India. Leave the brand names in the prompt and the model does not evaluate the minutes — it retrieves what it has read about the logos. You would be measuring brand recognition and calling it accuracy.
Take the names off and the model has nothing to go on except the two documents in front of it. That is the only condition under which the test means anything. It is also the condition under which we are most likely to lose, which is rather the point.
What the judge actually scores
Seven things, because these are the seven ways a set of minutes fails in front of the people who were in the meeting:
- Transcript fidelity — words correct, and critically, nothing invented. Confident fabrication is worse than a gap.
- Speaker separation — is each line attached to the right person, or are two people fused into one?
- Names, numbers and dates — the parts that get forwarded, and the parts a general model most often smooths over.
- Language handling — what survives a sentence that starts in one language and ends in another.
- Decisions captured — every decision the meeting reached, and no decision it did not.
- Action items — with a named owner and a date, or it is a sentence, not a task.
- Sendable as-is — could you forward it, unedited, to everyone who was in the room?
The prompt explicitly permits a tie. A rubric that cannot produce a tie is a rigged rubric.
What we are and are not claiming
We are claiming one thing: for a meeting held in a room — people around a table, a phone on the table, more than one language in the air — we think MeetNotes produces minutes you can send without editing more often than anything else you can install on a phone.
We are not claiming we beat a dedicated microphone worn on a shirt for all-day capture. We are not claiming we transcribe a Zoom call better than a bot sitting inside the Zoom call. Those are different tools doing different jobs, and pretending otherwise is how comparison pages lose their credibility in the first paragraph. There is an honest map of the whole category, including what each rival is genuinely better at, on the challenge page.
What happens when we lose
It gets published on the challenge page: which criterion we lost, by how much, and the date. Then we fix what lost and write down what changed.
That is not generosity, it is self-interest. A comparison we lose to a stranger's recording costs us one page of pride. The same failure discovered by a customer costs us the customer, silently, and we never find out why. We would very much rather have the page.
Send both outputs to support@getmeetnotes.com. The audio is optional — everything in the protocol works without sending us anything at all, and the 300 minutes are paid the same whether you send it or not, and whether we won or lost.
Start with your next meeting
You need a real meeting and about thirty minutes. MeetNotes is free to start with 100 minutes of recording included, so the test costs you nothing but the half hour. If you want the background first, the full rules and the judge prompt live here, and there is an honest breakdown of how the main meeting apps differ if you have not picked an opponent yet.
FAQ
Which AI should judge the comparison?
Whichever one you already trust — ChatGPT, Claude, Gemini or Copilot all work with the same prompt. Running it through two of them is better than one, because agreement tells you the gap is real and disagreement tells you the two outputs are close. What matters is that the judge is a model neither company controls.
Isn't a challenge run by the company that benefits from it still marketing?
Yes, and we would rather say so plainly than pretend otherwise. What makes it more than marketing is that every part of it is checkable by you: the judge prompt is published in full and names no product, the model doing the judging is one we have no access to, and we have committed to publishing the results we lose alongside the ones we win. If any of those three stopped being true, you should stop believing the page.
Which apps can I test MeetNotes against?
Any of them — Otter, Fireflies, Notta, Fathom, tl;dv, Read, Granola, Rev, Sonix, Descript, a Plaud device, a general AI chatbot, or your company's own internal tool. There is no exclusion list, because an exclusion list is how a company tells you what it is afraid of.
