The first result from our open challenge is in. One real ERP review call — several speakers, English with Hinglish running through it, dense technical vocabulary — put through Fireflies and MeetNotes from the identical audio file. Product names stripped from both outputs, then scored blind by general-purpose AI assistants that had no way of knowing which was which.
MeetNotes won, 8.7–9.2 against 4.7–5.2. Below is the whole scorecard, and then the part most comparison posts leave out: what a result like this does not prove.
The scorecard
| Criterion | Fireflies | MeetNotes |
|---|---|---|
| Overall evaluation | 4.7–5.2 | 8.7–9.2 |
| Transcript quality | ~4.0 | ~9.0 |
| Hinglish / multilingual | ~4.0 | ~9.5 |
| Technical terminology | ~3.5 | ~9.0 |
| Summary quality | ~6.5 | ~9.0 |
| Decisions captured | ~3.0 | ~9.0 |
| Action items | ~5.0 | ~8.0 |
| Factual caution | ~5.0 | ~8.5 |
The ranges are the spread between judges, not an error bar we calculated. The meeting is published anonymously — no participant, no company, and not one line of what was said. None of that was needed to score it, so none of it is here.
One disclosure about the rubric: this run predates the seven-criterion scorecard now published on the challenge page and used its own eight criteria, which is why the row names differ. Where a later result uses a scorecard other than the published seven, it says so too. We would rather disclose that than quietly relabel the rows to match.
What actually separated them
Look at where the gap is widest, because it is not random. The three worst rows for the losing side are technical terminology (~3.5), Hinglish (~4.0) and decisions captured (~3.0) — and those three are the same failure wearing different hats.
When a transcription model meets a word it does not expect, it does not stop. It substitutes the nearest word it does know. Do that to an ERP module name and you get a plausible English word in place of the thing the meeting was about. Do it to a Hindi verb in the middle of an English sentence and you get a phonetic mangle, or the clause quietly disappears. Then the summarising step reads that damaged transcript and writes a confident summary of a conversation that did not happen — which is exactly how you end up scoring 6.5 on summary quality while scoring 3.0 on decisions. The prose is fine. The decisions are gone, because the sentences carrying them were the ones that broke.
This is the failure we designed around, so it is the one where a gap should show up. Which brings us to the caveats.
What this does not prove
One meeting is one meeting. This is an N of 1. It shows what happened to this recording, on this day, with these judges. It is not a general ranking, not an average over many meetings, and not a prediction about your audio.
We picked the terrain, and it is our home terrain. An in-room, multi-speaker, two-language, jargon-heavy recording is the exact case MeetNotes was built for and it is nobody's easiest problem. Fireflies is built as a meeting bot: it joins your Zoom or Teams call and receives clean per-participant audio directly from the conferencing platform. That is a genuinely easier signal to work with, and Fireflies is good at it. Put a Fireflies bot in an English video call and this comparison would look very different — quite possibly the other way round.
We ran it. We are the interested party. We stripped the names before scoring and used judges we do not control, which is the best correction we know of, but it is a correction, not a proof of neutrality.
If that all sounds like we are talking ourselves down, we are — deliberately. A comparison page that oversells one result is worth nothing the first time somebody checks it.
Come and beat it
The whole point of publishing the method is that you do not have to believe any of the above. Send us one real meeting's audio and whatever your current tool made of it, and you get back:
- our transcript and our minutes, as a PDF and as Markdown, so you can paste them into any AI without a converter in the way;
- our blind comparison of the two, with the judging assistant named so you can repeat it;
- 200 free minutes on the MeetNotes account for the email you wrote from.
You then decide whether we may publish it. Not before — after you have seen it. Silence is a no, the minutes are paid either way, and if your own run of the judge prompt disagrees with ours, we want that one most of all. Losses go on the challenge page next to the wins, with the criterion we lost on.
Full rules, the conditions that make an entry usable, and the judge prompt →
If you would rather not send us anything at all, don't — every part of the method works without us. And if you are still choosing which tools to line up, here is the honest map of the category, including what each one is genuinely better at than we are.
FAQ
Is Fireflies bad at transcription?
No — and this result should not be read that way. Fireflies is built around meeting bots that join video calls, where it receives a clean separate audio stream per participant and performs well. This test used in-room audio from a single microphone, with two languages mixed inside sentences and heavy domain jargon, which is the hardest end of the problem and is not the case Fireflies is designed around. A different meeting would produce a different scorecard, which is precisely why we tell you to run your own.
How do I know the comparison was fair?
You do not have to take it on faith, which is the design. The product names were removed from both outputs before any AI scored them, so the judges could not favour a brand they had read about. The rubric is published in full on the challenge page and permits a tie. And every entrant gets our raw outputs in Markdown, so anyone can re-run the scoring themselves in whichever assistant they trust and tell us we got it wrong.
Why is the meeting anonymous?
Because nothing about the participants was needed to score it. We publish the scorecard, the judge's reasoning and the shape of the meeting — length, languages, how technical — and never the contents, the people, the company or the audio. That rule holds for every result on the page, including ones we lose.
