MeetNotes blog

How to Actually Test Whether a Transcription App Is Accurate (in 30 Minutes, on Your Own Audio)

24 Aug 2026

The fastest honest test takes thirty minutes: transcribe two minutes of your own real meeting audio by hand, count the words the app got wrong, and divide by the total. That number — the word error rate — is the only accuracy figure about your meetings that anyone can defend. This post explains how to do it properly, and then why the percentage is still not the thing that decides which app you should use.

Why the vendor's 99% is not a lie, and still not useful

When a transcription company publishes an accuracy figure, it is usually real. It was measured on a standard benchmark: clean recordings, one speaker at a time, close microphone, mostly American English, mostly prepared speech. On that material, modern speech models genuinely do reach the high nineties.

Your meeting is not that material. Your meeting has a fan, a road outside, a speakerphone, six people, two of whom interrupt, an accent the benchmark never contained, and at least one sentence that starts in one language and finishes in another. Accuracy on conversational multi-speaker room audio is a completely different problem from accuracy on a benchmark, and every published number lands somewhere between "optimistic" and "irrelevant" once it meets a real room.

This is not a scandal and nobody is being dishonest. It just means the only measurement that answers your question is one made on your audio.

Word error rate, the practical version

WER is the industry's standard metric, and you can compute a usable version of it in twenty minutes with no tools.

Step 1 — pick two minutes. Take a real meeting recording. Choose a two-minute stretch that is representative, not the clearest part. Include a bit of cross-talk if your meetings have cross-talk.

Step 2 — write the truth. Transcribe those two minutes by hand, exactly as spoken. This is the tedious part and there is no way around it: without a reference transcript there is nothing to measure against. Two minutes of conversation is roughly 300 words and takes about fifteen minutes to type.

Step 3 — count three kinds of error against the app's output:

  • Substitutions — a wrong word ("Kapil" → "capital")
  • Deletions — a word that was spoken and is missing
  • Insertions — a word that appears but was never said

Step 4 — divide.

WER = (substitutions + deletions + insertions) / total words in your reference

300 reference words with 21 errors is a 7% WER, or 93% accuracy. Do this for each app on the same two minutes and you have a real, comparable number.

What is good? On clean single-speaker audio, under 5% is normal. On genuine multi-speaker room audio, anything under 10% is strong and under 15% is usable. If you see 25% or more, the transcript will cost you more time to fix than it saved.

Two rules that keep the test honest: use the same audio for every app, and do not re-record because a result was bad. A second attempt tests the meeting, not the tool.

The five failures that matter more than the percentage

Here is the uncomfortable part. Two apps can post an identical WER and be worlds apart in usefulness, because WER weights every word the same and your meeting does not.

1. Invented text. A deletion leaves a visible gap you will notice. A confident fabrication reads perfectly and is wrong, and it is the failure most likely to survive into a document you send to twelve people. WER counts these the same. You should not. Read the output specifically hunting for sentences that sound plausible and never happened.

2. Names, numbers and dates. "Friday" becoming "Thursday" is one substitution out of 300 — 0.3% of your WER, and 100% of the reason someone missed a deadline. Score proper nouns, amounts and dates separately, and weight them heavily.

3. Speaker attribution. WER does not measure it at all. A perfect transcript with two people fused into one speaker is useless for minutes, because you cannot tell who committed to what. Count how often two speakers merge and how often one person is split in two.

4. Mixed-language sentences. "Toh basically we'll ship the Q3 numbers by Friday" is one sentence in two languages and it is how a very large part of the world talks in meetings. Watch what each tool does with it: some translate it, some drop it, some spell the Hindi phonetically in English letters. No pricing page tells you which, and no language count predicts it.

5. Whether it survives into the minutes. You do not read transcripts, you read minutes. A tool can transcribe beautifully and still produce a summary that misses the one decision the meeting reached. Test the output you will actually send.

The blind A/B test (the thirty-minute version)

If you do not want to hand-transcribe anything, there is a shortcut that measures the thing you actually care about — which output is more usable — rather than the word count.

  1. Run the same recording through both apps. One attempt each.
  2. Delete every product name, logo and footer from both outputs. Label them A and B.
  3. Paste both into ChatGPT, Claude or Gemini and ask it to score each on transcript fidelity, speaker separation, names and numbers, language handling, decisions captured, action items with owners, and whether it is sendable unedited.
  4. Ask it to name a winner per criterion, an overall winner, and anything either one invented.

Step 2 is the step people skip and the one that decides whether the result means anything. Large language models have read years of marketing about the big incumbent brands and very little about small ones. Leave the names in and the model recalls its impressions of the logos instead of reading the documents; you will be measuring brand recognition and calling it accuracy.

A ready-made version of that rubric — blind, naming no product, and explicitly permitting a tie — is published in full on our open challenge page. Copy it, use it on anything, including against us.

Put it to work

We publish that prompt because we would rather be tested than believed. Run it on MeetNotes against whatever you use today: we publish the head-to-heads we lose along with what we fixed, and we send 300 free minutes to anyone who sends one in, win or lose. The rules are here.

If you are still choosing which apps to put in the test, there is an honest map of the category in Otter vs Fireflies vs Notta vs Plaud vs MeetNotes — including what each one is genuinely better at than we are.

FAQ

What is a good word error rate for meeting transcription?

On clean, single-speaker, close-microphone audio, under 5% is the normal expectation from a modern speech model. On real multi-speaker room audio — several people, background noise, cross-talk — under 10% is strong and under 15% is usable. Above roughly 25%, correcting the transcript costs more time than writing notes by hand would have. Always compare numbers measured on the same recording; a WER from one dataset says nothing about performance on another.

How do I calculate word error rate without any tools?

Hand-transcribe two representative minutes of your own audio as a reference (roughly 300 words, about fifteen minutes of typing). Then count substitutions, deletions and insertions in the app's version, add them up, and divide by the number of words in your reference. That is your WER. Use the identical two minutes for every app you are comparing, or the numbers are not comparable.

Why shouldn't I just trust published accuracy percentages?

Because they are almost always measured on benchmark corpora — clean, single-speaker, prepared speech in a small number of accents — and your meetings are none of those things. The figures are usually accurate for what they measured; they simply did not measure your room. A two-minute test on your own audio beats every published percentage, because it is the only one taken under the conditions you will actually use the tool in.

Your next meeting can write its own minutes

Record on your phone, get a speaker-labelled transcript and minutes as PDF or Word. Minutes in 100+ languages. Free to start with 100 minutes included.

Download on the App StoreGet it on Google Play