When the Transcript Is Wrong

The hidden first layer of AI meeting governance

AI meeting governance often focuses on whether an AI-generated summary, decision or action item is accurate.
But there is an earlier question:
Was the transcript itself accurate?
Meeting transcripts are increasingly used as source material from which AI systems generate summaries, identify decisions, extract commitments and create follow-up information.
If something important is incorrectly transcribed, downstream AI can interpret the incorrect text perfectly — and still produce the wrong business record.

The transcript is not the meeting

A Microsoft Teams .vtt transcript may look like a factual record of a meeting.
But it is a machine-generated representation of what was spoken.

Microsoft itself recognizes that language settings affect transcription and caption accuracy. Its Teams documentation advises users to select the language actually being spoken and, where multilingual speech recognition is used, for participants to set their correct spoken language for transcription and caption accuracy.
That does not mean Teams transcription is generally unreliable.
It means the transcript should not automatically be treated as an infallible record of what was said.

What does the evidence tell us about accents?

This distinction is important.
There is independent research showing that automatic speech recognition (ASR) systems can have greater difficulty with accented speech and particularly with accents underrepresented in their training data.
For example, research published at EMNLP 2023 by Prabhu and colleagues describes speech accents as a significant challenge for state-of-the-art ASR systems and examines recognition performance for underrepresented accents.
That research did not test Microsoft Teams specifically.
It therefore should not be used to claim that Teams is less accurate for particular nationalities, countries or accents.
What it does establish is a broader technology risk relevant to organizations using automated transcription: speech recognition accuracy can vary with the characteristics of the speech being processed.

Multilingual meetings add another consideration

Microsoft Teams supports multilingual speech recognition and numerous spoken languages and regional variants.
Microsoft's own documentation emphasizes the importance of correctly identifying the language being spoken for transcription and caption accuracy.
This matters in ordinary international business meetings, where participants may:

  • speak different first languages;
  • speak English with different accents;
  • use more than one language during a meeting;
  • use specialized company or technical terminology.

The governance conclusion should therefore be carefully stated.
Multilingual or accented speech does not mean that a Teams transcript will be inaccurate.
Rather organizations should recognize that automated transcription is a generated representation of speech and that important information may warrant verification before the transcript is relied upon as an authoritative business record.

A small transcription error can become a large governance error

Consider a hypothetical example.
What was said:
"We haven't approved the contract.”
What the transcript records:
"We have approved the contract.”
Only one word has changed.
But the meaning has reversed.
If an AI system subsequently processes the incorrect transcript, it could reasonably conclude:
Decision: The contract was approved.
The downstream AI has not necessarily misunderstood its source.
Its source was wrong.
This example is illustrative; it is not presented as an observed Microsoft Teams error.

What deserves verification?

This does not mean organizations need people to compare every transcript word-by-word with a meeting recording.
That would defeat much of the productivity benefit of AI.
A more practical governance approach is to identify information where an error could materially change the business record.
Examples include:
Decisions
Was something actually approved, rejected or deferred?
Commitments
Did the identified person actually agree to perform the action?
Dates and deadlines
Was the commitment Friday, next Friday or the end of the month?
Amounts and numbers
Was the amount, percentage or quantity captured correctly?
Names and responsibilities
Was the correct person, organization or responsibility identified?
Meaning-changing words
Could a missing or incorrectly inserted word such as “not”, “don't”, “can't” or “haven't” reverse the meaning?These are small differences capable of producing materially different business outcomes.

Accuracy percentage is not the governance question

A transcript could be highly accurate overall and still contain one material error.
Imagine a long meeting in which thousands of words are transcribed correctly, but one sentence incorrectly records whether a contract was approved.
The overall transcription accuracy might appear excellent.
From a governance perspective, however, the one incorrect sentence may be more important than all the correctly transcribed conversation around it.
This suggests that the appropriate question is not simply:
“How accurate is the transcript?”
It is:
“Is the information we are about to rely upon accurate?”

Who verified the source?

AI governance increasingly addresses the reliability of AI-generated outputs.
Meeting governance introduces an additional dependency:
AI-generated outputs may themselves be based on another machine-generated output — the transcript.
That creates a chain of reliance.
Speech → Transcript → AI interpretation → Business record
Good governance therefore needs to consider not only whether the final AI output appears reasonable, but whether important information remains supported by reliable evidence throughout that chain.
Which brings us back to the fundamental question:

Who Verified This?

Sources and further reading

Microsoft Teams documentation advises users to select the language actually being spoken and provides multilingual speech-recognition settings intended to improve transcription and caption accuracy.
Microsoft states that AI-generated content in Teams Intelligent Recap is based on the meeting transcript and cautions that AI-generated content may be inaccurate, incomplete or inappropriate.
Independent research into automatic speech recognition has documented challenges associated with accented speech, including performance degradation involving underrepresented accents. This research concerns ASR technology generally and should not be interpreted as a measurement of Microsoft Teams transcription accuracy.
Microsoft: Meeting options in Microsoft Teams — multilingual speech recognition and transcription accuracy
Microsoft: Use live captions in Microsoft Teams meetings — spoken-language selection and caption accuracy
Microsoft: Recap in Microsoft Teams — transcript-based AI content and AI accuracy limitations
Independent research: Prabhu, D., Jyothi, P., Ganapathy, S. & Unni, V. (2023), Accented Speech Recognition With Accent-specific Codebooks, Proceedings of EMNLP 2023, pp. 7175–7188.