The second result from our open challenge is in, and it is a more interesting one than the first — because this time the other side won a criterion, and won it on the thing most people think they are buying.
A 38-minute cross-border partnership call. Two sides, commercial and technical vocabulary throughout. Teams transcribed the call it was hosting; MeetNotes worked from a recording of the same conversation. Both outputs then had every product name stripped and were scored blind by general-purpose AI assistants with no way of knowing which was which.
Microsoft Teams ~6.8. MeetNotes ~8.8. Here is the whole table, and then the five words that decided it.
The scorecard
| Criterion | Microsoft Teams | MeetNotes |
|---|---|---|
| Overall, transcript against transcript | ~6.8 | ~8.8 |
| Transcript completeness | 8.5 | 8.2 |
| Readability | 6.5 | 9.0 |
| Speaker separation | 7.5 | 8.5 |
| Names and product terminology | 5.0 | 9.0 |
| Numbers and commercial detail | 8.0 | 8.8 |
| Context understood | 6.0 | 9.2 |
| Factual caution | 6.5 | 8.8 |
| Concision | 5.0 | 9.2 |
| Business usefulness | 5.8 | 9.0 |
Read the second row first. Teams beat us on completeness, 8.5 to 8.2. It caught more of the words — more of the crosstalk, more of the small talk, more of the half-sentences that go nowhere. If your definition of an accurate transcript is "how much of the audio ended up as text", Teams won this test and we lost it.
We are publishing that row at the top because it is the one a vendor would bury, and because it is the whole argument. A transcript is not scored by weight.
The five words
Here is what Teams did with the proper nouns the meeting was actually about.
| Teams heard | The word was |
|---|---|
| "Zero" | Xero |
| "my orb" / "my op" / "my ops" | MYOB |
| "Israel" | New Zealand |
| "Datablix", "Jenny", "Vortex AI" | Databricks, Genie, Cortex AI |
There is a fifth we are not printing: the name of the product the entire call was about, rendered throughout as an unrelated personal name. Printing it would identify the company whose meeting this was, and that is theirs to give, not ours to take. Everything else here is a generic product name that identifies nobody.
Now read them as a person would a week later. Xero and MYOB are the two accounting platforms the market in question runs on — the transcript says the meeting discussed nothing and something called "my orb". The entire conversation was about entering New Zealand; the transcript relocates it to Israel, and later renders "Indian time difference and New Zealand time difference" as "Indian time difference and user time difference". The technical comparison — Databricks has Genie, Snowflake has Cortex AI — survives as a sentence about Datablix, Jenny and Vortex AI.
Every one of those lines is complete. Not one word is missing. They are simply about different companies in a different country.
Why a more complete transcript scored lower
This is the mechanism, and it is worth understanding whichever tool you end up using.
When a speech model meets a word it does not expect, it does not leave a gap. It substitutes the nearest thing in its vocabulary and moves on, at full confidence. "Xero" becomes "zero" because "zero" is a word and "Xero" is a company. "MYOB" becomes "my orb" because letters spoken quickly sound like a phrase. The output is fluent, complete, and wrong — and it is wrong precisely at the proper nouns, which are the only part of a business transcript anybody forwards.
That is why completeness and usefulness came apart by three full points here. The words that broke were not random words. They were the ones carrying the meaning.
What this does not prove
One meeting is one meeting. An N of 1. It is what happened to this recording, on this day, with these judges. Not a general ranking, not an average, not a prediction about your audio.
This was not our home terrain — it was theirs. Our first published result was an in-room Hinglish meeting, which is the case we were built for and nobody's easy problem. This one is the opposite: an online call, on Teams, transcribed by Teams from the call it was hosting. That is the clean, direct signal that the whole meeting-bot category exists to exploit, and it is a harder comparison for us to win, not an easier one.
We only scored the transcript. The Teams export we were given contained the transcript and nothing else — no summary, no action items. So there was nothing on that side to score our minutes against, and we did not invent a score for it. That is a fact about the file we received, not a claim that Teams cannot produce a recap. It can, and for many teams it does.
We ran it. We are the interested party. Stripping the names before scoring and using judges we do not control is the best correction we know of. It is a correction, not a proof of neutrality.
When Teams is still the right answer
If your meeting is a Teams call inside a Microsoft company, a lot of this is beside the point, and we have written the honest version of that comparison. Teams already has the audio. Every line already carries a speaker name and a timestamp, because the call already knows who joined it. The transcript lands in the organiser's own OneDrive under retention rules your IT department already set — a category we do not compete in at all. And Microsoft lists more than forty transcription languages, with live translated transcription on its Premium tier.
The gap this test found is narrower than the scores make it look: it is what happens to unfamiliar proper nouns, and to a sentence that changes language halfway through. Worth knowing that Teams transcription runs on one selected spoken language per meeting — if the room speaks another, Microsoft's own help page says it prompts the organiser to change it, and the change applies to everyone. That is a sensible design for a call held in one language. It is not a design for a room that switches mid-sentence.
Check it yourself in twenty minutes
You do not have to believe a table on a vendor's website, and you should not. The test is four steps and you can run it on a meeting you already have:
- Take one real meeting — twenty minutes is plenty, with real interruptions and real accents.
- Put the same audio through both tools. No second attempts, no picking the better run.
- Delete every product name from both outputs. Label them A and B. If the judge can tell whose is whose, you have measured brand recognition, not transcription.
- Paste both into ChatGPT, Claude or Gemini with our judge prompt — a prompt written to be winnable by the other side, which invites a tie.
Then check the proper nouns by ear. That is the row that matters, and it is the one you can verify without trusting anybody.
Send us the result and you get 200 free minutes on the account for the email you wrote from, our transcript and minutes as PDF and Markdown, and our blind comparison — win or lose, published or not. Losses go on the challenge page beside the wins, with the criterion we lost on.
The full protocol and the judge prompt →
FAQ
Is Microsoft Teams transcription accurate?
In this test it was accurate on volume and weak on vocabulary. Teams scored 8.5 out of 10 for completeness — it captured more of the words than we did — but 5.0 on names and product terminology, because it rendered Xero as "Zero", MYOB as "my orb", New Zealand as "Israel", and Databricks, Genie and Cortex AI as "Datablix", "Jenny" and "Vortex AI". For a call in one language with common vocabulary, Teams transcription is solid. The failures cluster on proper nouns, unfamiliar product names, and sentences that change language mid-way.
Why does Teams get product names and proper nouns wrong?
Because a speech model does not leave a gap when it hears a word it does not know — it substitutes the nearest word in its vocabulary and carries on at full confidence. "Xero" is not in most vocabularies; "zero" is. "MYOB" spoken quickly sounds like a phrase. The result is a transcript that is complete and fluent and wrong, and wrong specifically at the words that carry the meaning.
How do I improve Microsoft Teams transcription accuracy?
Three things help, and none of them need another product. Set the meeting's spoken language to what the room is actually speaking before you start, because Teams transcribes against one selected language. Type unusual product names, company names and acronyms into the meeting chat, so the reader of the transcript has them spelled correctly somewhere. And read the proper nouns in the transcript against your own memory before you forward it — those are the lines that fail, and they fail silently.
Can Microsoft Teams transcribe an in-person meeting?
Not on its own. Teams transcription works on a Teams call — it transcribes the stream of a meeting it is hosting. A meeting held around a table with no call running has nothing for it to transcribe. For that you need something recording the room itself, which is what MeetNotes does from a phone.
Does Microsoft Teams transcription support more than one language in a meeting?
Teams lists more than forty transcription languages and offers live translated transcription on its Premium tier, but a meeting transcribes against one selected spoken language. Microsoft's own help page says that if people speak a different language, Teams prompts the organiser to update the setting — and changing it changes the transcript and caption language for everyone. That works for a call held in one language. It is not built for a sentence that starts in Hindi and finishes in English.
Does Microsoft Teams produce meeting minutes and action items?
It can — an AI recap of a Teams meeting is a Microsoft feature and this test says nothing about it. The export we were given for this comparison contained the transcript only, with no summary or action-item section in the file, so there was nothing on that side to score. We scored transcript against transcript and left it there.
Which is more accurate for meeting notes, Teams or MeetNotes?
On this one recording, blind-scored, MeetNotes came out ~8.8 to Teams' ~6.8 overall — but Teams won completeness, and it was a Teams call, which is Teams' strongest case. One meeting is one meeting. The only comparison that decides anything for you is the one you run on your own audio, and the protocol and judge prompt are published in full so that you can.
