Open challenge

Don't take our
word for it.
Take your audio.

Every "best AI note taker of 2026" list you have read was written by someone who sells one of the entries. So we are not going to tell you we are the best. We are going to hand you a test, a judge prompt that names nobody, and an open invitation to run it against anything you like — and we will publish the ones we lose.

Get the judge prompt See the first result

The claim

Here is exactly what
we are claiming.

What we say

For a meeting held in a room — people around a table, one phone on the table, more than one language in the air — we believe MeetNotes produces minutes you can send without editing more often than anything else you can install on a phone today.

What that is worth

Nothing, until you check it. It is a belief we hold because of what we chose to build, and beliefs held by founders about their own products are the least reliable information on the internet. Which is why the rest of this page is a test rather than an argument.

Note what we are not claiming: that we beat a dedicated microphone worn on a shirt, that we transcribe a Zoom call better than a bot sitting inside the Zoom call, or that we are cheaper than free. Different tools, different jobs. This challenge is about the room.

The rules

Four steps.
About thirty minutes.

01

Record one real meeting

Not a podcast, not a scripted demo — a real meeting with real interruptions, real accents, and people talking over each other. Twenty minutes is plenty. Keep the audio file.

02

Run it through both

Put the same file through MeetNotes and through whatever you use today, or whatever a listicle told you to use. Same audio, same day, no second attempts and no cherry-picking a good run.

03

Strip the names, label them A and B

Delete every logo, footer and product name from both outputs. If the judge can tell whose is whose, you have not run a test — you have run a brand-preference survey.

04

Let an AI you already trust judge it

Paste both into ChatGPT, Claude, Gemini, Copilot, or all of them, with the prompt on this page. You are not asking us who won. You are asking a machine with no stake in the answer.

One rule we ask you to hold us to: run it once. If you record the meeting again because the first result was bad, you are testing the meeting, not the tool. Whatever came out of the first run is the result — for us as much as for them.

The judge prompt

Copy this into
any AI you trust.

It is blind, it names no product, and it explicitly permits a tie. Read it before you use it — if you can find a way it favours us, tell us and we will change it.

You are an impartial evaluator. Below are two sets of AI-generated
meeting minutes, A and B, produced from the SAME recording of the same
meeting. The product names have been removed. You do not know which is which,
and you must not guess or speculate about which brand produced either one.

Score A and B independently, 0-10, on each of these seven criteria:

1. Transcript fidelity - words correct, nothing invented.
2. Speaker separation - each line attributed to the right person.
3. Names, numbers and dates - proper nouns, amounts and deadlines correct.
4. Language handling - sentences that mix two languages kept intact,
   not translated away, dropped, or spelled phonetically.
5. Decisions captured - every decision the meeting reached is present,
   and no decision is present that the meeting did not reach.
6. Action items - each one carries a named owner and a due date.
7. Sendable as-is - could this be forwarded to the whole room unedited?

Then output:
- a table of the 14 scores,
- the winner of each individual criterion (or "tie"),
- an overall winner (or "tie", if it is genuinely a tie),
- the single strongest reason the loser lost,
- anything either one INVENTED that was not in the meeting.

Be blunt. Do not be diplomatic, do not hedge, and do not try to find
something nice to say about both. If one is clearly better, say so plainly.

--- A ---
[paste the first set of minutes here]

--- B ---
[paste the second set of minutes here]

Works in ChatGPT, Claude, Gemini and Copilot. Running it in two of them and comparing the verdicts is better than running it in one.

The scorecard

Seven things minutes
are actually judged on

Not features. Not language counts on a pricing page. These are the seven ways a set of minutes fails in front of the people who were in the meeting.

CriterionWhat to look for
Transcript fidelityRead along with the recording. How many words are wrong, missing, or invented? Invented text is the worst failure — it is confidently wrong.
Speaker separationIs every line attached to the right person? Count how often two people are fused into one speaker, or one person is split into two.
Names, numbers and datesProper nouns, amounts, quantities, and deadlines. These are the parts of a transcript that get forwarded, and the parts a general-purpose model most often smooths over.
Language handlingWhat happens to a sentence that starts in one language and finishes in another? Most tools quietly translate it, drop it, or transcribe it phonetically.
Decisions capturedList every decision the meeting actually reached. Now check which ones survived into each set of minutes, and whether either invented a decision nobody made.
Action items with ownersAn action item without a named owner and a date is a sentence, not a task. Score only the ones that carry both.
Sendable as-isCould you forward the output to everyone who was in the room, unedited, without embarrassment? This is the only question that matters at 7pm.
The field

Compare us against
any of these.

There is no exclusion list, because an exclusion list tells you what a company is afraid of. Here is the honest map of the category, including what each kind of tool is genuinely better at than we are.

Meeting bots

Otter, Fireflies, Fathom, tl;dv, Read, Granola

Built for calls: a bot joins your Zoom, Teams or Meet and transcribes the stream. Excellent at that job. There is no bot to send into a room with a table in it.

Hardware recorders

Plaud Note, Plaud NotePin, and similar

A dedicated microphone plus software, roughly ₹17,000–21,000 in India, typically with a subscription on top. Genuinely better than a phone for all-day wearable capture and phone-call pickup.

Transcription services

Rev, Sonix, Descript, Notta

Strong transcripts, editor-first workflows, and in some cases human transcribers. Built around producing a document from a file rather than around a meeting that ends at 6pm and needs sending at 6:10.

General AI chatbots

ChatGPT, Claude, Gemini

Superb summarisers of text you hand them. They do not record your meeting, do not separate speakers, and make you repeat the whole ritual by hand after every single one.

Why we will take the bet

Five decisions that
show up in the output

The room, not the call

MeetNotes was built for a phone lying on a table in a room with eight people in it — the case the meeting-bot category structurally cannot serve. That is the whole design, not a mode we bolted on.

Sentences that switch language

There is no language to select before recording. A sentence that starts in Hindi and ends in English is the normal case here, not an unsupported edge — and it is the single fastest way to break a transcription tool.

Minutes in 100+ languages

The room speaks whatever it speaks; the minutes come out in any of 110 languages. Record in Hinglish, send the MoM in English. For comparison, Otter's own help centre lists six transcription languages.

Speaker names you confirm

The app asks you to confirm who is who before you export, instead of silently guessing and letting a wrong name go out to twelve people. Guessed names are the failure nobody notices until it is embarrassing.

Minutes shaped like minutes

Summary, key points, decisions, and action items with owners and dates — the format an MoM has had for a century — as a PDF and an editable Word file, ready for WhatsApp. Not a wall of bullet points.

How to enter

You send two things.
We send four back.

This is an exchange, not an inbox. We are not asking you to post an opinion — we are asking for one recording and the output your current tool made from it, and you get our full working in return, in formats you can re-judge yourself in any AI you like.

You send

The audio, and their output

  • The audio file. The recording itself, so we run it through MeetNotes ourselves rather than taking your word for what we would have produced.
  • What your current tool made from that same audio — its transcript and its summary. Two files, or one export containing both; whatever their export button gives you.
You get back

Our working, and the verdict

  • Our transcript and our minutes, as a PDF and as Markdown — Markdown specifically so you can paste it straight into any AI without a converter in the way.
  • Our blind comparison: both outputs stripped of product names, scored by a general-purpose AI assistant, with the assistant named so you can repeat it.
  • 200 free minutes, on the MeetNotes account for the email you wrote from. Sign in to the app with that same email and they are already there.

You do not have to accept our verdict, and we would rather you did not. You have both sets of minutes in Markdown — run the judge prompt yourself, in whichever assistant you trust, as many times as you like. If your run disagrees with ours, send it to us: that is a more interesting result than agreement, and it is paid the same.

What makes an entry usable

Five conditions, and
they bind us too

A real meeting. More than one person, actually talking to each other. A scripted demo or a podcast tells us nothing, because neither has the interruptions and cross-talk that break these tools.

Ten minutes or more. Short clips flatter everyone. The failures that matter — a speaker drifting onto the wrong label, a decision quietly dropped — need length to show up.

The same audio, both sides. Not the same meeting recorded twice on two devices. The identical file, or the comparison measures microphones instead of software.

One attempt each. Whatever came out of the first run stands. Re-running until a result improves tests the meeting, not the tool — and that rule costs us more than it costs you.

You have the right to share the recording. This is the one we cannot check and will not guess at. If your meeting was confidential, or the people in it would not expect a copy to leave your company, do not send it — run the whole test yourself instead. Everything on this page works without us ever seeing your audio; you only lose our half of the working.

Send it to support@getmeetnotes.com. A person reads every entry — at this volume that beats a pipeline, and it means nothing you send is processed by anything you did not expect.

Download on the App StoreGet it on Google Play

Get the app first if you have not already — a recording of your own is the only thing the test really needs.

What we publish

Your meeting stays
your meeting.

We never publish
  • Anything that was said in your meeting — no transcript, no summary, no quotes, not one line.
  • Who was in the room, their names, or your company's name, unless you tell us in writing that we may.
  • Your audio. Ever, to anyone, for any reason.
We do publish, with your go-ahead
  • The scorecard, and the judge's reasoning for it.
  • A one-line description of the kind of meeting it was — length, languages, how technical — because a score means nothing without it.
  • Which tool it was against, named.

You decide after you have seen it, not before. Sending an entry gives us no publishing rights at all. We do the comparison, send you everything, and then ask — and you can say yes, yes-but-anonymously, or no. Silence is a no. A yes covers the scorecard and lets us use it in our marketing; it never covers your meeting's contents, its participants, or your audio, and you can withdraw it later and we take the result down. Your 200 minutes do not depend on the answer. See privacy and delete my data.

Results

Every head-to-head,
including the losses.

26 August 2026

MeetNotes vs Microsoft Teams

Winner: MeetNotes — except on raw completeness, which Teams won

A 38-minute cross-border partnership call: two sides, commercial and technical vocabulary throughout, and a run of product names that decide whether the notes still mean anything a week later. Published anonymously — no participant, company or meeting content is identified beyond the individual misheard terms below, and none was needed to score it.

Judged by: Run blind — the product names were stripped from both outputs before scoring, so neither could be traced to a brand. Scored by more than one general-purpose AI assistant.

Rubric: Transcript against transcript. The supplied Teams export contained the transcript only, with no summary or action-item section in the file, so there was nothing on that side to score our minutes against — that is what was in the export we were given, not a statement about what Teams can produce. The nine criteria scored are listed as scored; like the run below, they are not the seven published above.

CriterionMicrosoft TeamsMeetNotes
Overall, transcript against transcript~6.8~8.8
Transcript completeness8.58.2
Readability6.59.0
Speaker separation7.58.5
Names and product terminology5.09.0
Numbers and commercial detail8.08.8
Context understood6.09.2
Factual caution6.58.8
Concision5.09.2
Business usefulness5.89.0

Teams won completeness: it caught more of the words, including small talk we dropped. Then it spent that advantage on the proper nouns the meeting was actually about.

Microsoft Teams heard

“Zero”

The word was

Xero

Microsoft Teams heard

“my orb / my op / my ops”

The word was

MYOB

Microsoft Teams heard

“Israel”

The word was

New Zealand

Microsoft Teams heard

“Datablix · Jenny · Vortex AI”

The word was

Databricks · Genie · Cortex AI

One further error is not shown, because printing it would identify the company whose meeting this was: the name of the product the entire call was about, rendered throughout as an unrelated personal name.

The full write-up of this result

25 August 2026

MeetNotes vs Fireflies

Winner: MeetNotes

A long ERP review call — English with Hinglish stretches running through it, several speakers, and heavy domain-specific vocabulary. The kind of meeting where the words that carry the decisions are exactly the words a general-purpose model has never seen. Published anonymously: no participant, company or content is identified, and none was needed to score it.

Judged by: Run blind — the product names were stripped before scoring, so neither output could be traced to a brand. Scored by more than one general-purpose AI assistant; the ranges below are the spread between them, not a margin of error we calculated.

Rubric: This run predates the seven-criterion scorecard published above and used its own eight criteria, listed as scored. Where a run uses a scorecard other than the published seven, it says so here rather than quietly re-labelling the rows to match.

CriterionFirefliesMeetNotes
Overall evaluation4.7–5.28.7–9.2
Transcript quality~4.0~9.0
Hinglish / multilingual~4.0~9.5
Technical terminology~3.5~9.0
Summary quality~6.5~9.0
Decisions captured~3.0~9.0
Action items~5.0~8.0
Factual caution~5.0~8.5

The full write-up of this result

One meeting is one meeting. A result here shows what happened to this recording, on this day, with these judges — not a general ranking, not an average, and not a claim about how either tool performs on your audio. That is the entire reason the protocol is published: the only comparison that decides anything for you is the one you run yourself.

FAQ

Fair questions about
a challenge run by
an interested party

Is this a real challenge or a marketing gimmick?

It is real, and it is deliberately designed so we can lose it. The judge prompt on this page is blind and names no product, we do not see your audio unless you send it to us, and we have committed to publishing head-to-heads we lose on this page along with what we changed afterwards. A comparison you run yourself, judged by a model we do not control, is the only kind of proof worth anything.

Which apps can I compare MeetNotes against?

Any of them. Otter, Fireflies, Notta, Fathom, tl;dv, Read, Granola, Rev, Sonix, Descript, a Plaud device, a general AI chatbot, or your company's in-house tool. We have no exclusions, because an exclusion list is how you tell people what you are afraid of.

Which AI should judge it?

Whichever you already trust — ChatGPT, Claude, Gemini or Copilot. Running the same prompt through two or three of them is better, because it shows you whether the result is a real gap or one model's taste. If they disagree, that itself is the answer: the two outputs are close.

Why blind? Why strip the product names?

Because large language models have read the internet, and the internet is full of marketing about the big brands. Leave the names in and you are measuring brand recognition. Take them out and you are measuring the minutes.

What do I get for sending you my comparison?

200 free minutes, plus our transcript and minutes as a PDF and Markdown and our blind comparison of the two. The minutes land on the MeetNotes account for the email you wrote from — sign in to the app with that same email and they are waiting. Paid whether MeetNotes won or lost, and paid whether or not you let us publish it.

Can you just read my Otter or Fireflies share link?

No, and it is worth knowing why before you rely on one. A share link from those tools resolves to an app that loads its content only after checking who is signed in — to anyone who is not you it is a login wall, with no transcript, no summary and no audio behind it. We are also not going to build something that scrapes a competitor's product to get around that: it would breach their terms, it would break the first time they changed a page, and a challenge whose credibility rests on playing fair cannot be run by cheating at the intake step. Use their export button instead — it is your data and it takes one click.

What happens if MeetNotes loses?

We publish it here with the criterion we lost on, we fix what lost, and we tell you when it is fixed. We would rather find out from a stranger's recording than from a customer who quietly stopped using it.

Do you need my meeting audio?

No. The test is yours to run and yours to keep — everything on this page works without sending us anything. The audio only helps if you want us to diagnose a specific failure, and even then we will delete it on request. See the privacy policy for what we store and for how long.

Sources

Where the third-party claims on this page come from

  • Otter transcription languages (six: English, Spanish, French, German, Japanese, Chinese Simplified) — help.otter.ai — Supported languages, checked 24 August 2026
  • Plaud device pricing in India (₹17,000–21,000, typically plus subscription) — Indian retail listings, checked August 2026

Deliberately absent: competitor pricing. Prices change without notice, and a stale price on a comparison page is a false statement about someone else's business. Check theirs on their own site, and ours on our pricing page.

Run the test on your own meeting

Free to start — 100 minutes of recording included, no card. Then send us the result, win or lose, for 200 more.

Download on the App StoreGet it on Google Play