Est.

Accent and Dialect Effects on AI Debate Judge Scores

Speech-to-text bias determines debate outcomes before judges even score arguments.

Editorial team · · 11 min read
Cover illustration for “Accent and Dialect Effects on AI Debate Judge Scores”
Bias Auditing and Fairness · October 5, 2026 · 11 min read · 2,371 words

Accent and dialect bias in AI-judged debate does not begin with the judge. It begins earlier, in the system that turns a debater's voice into the words the judge will ever see, and that earlier stage decides the outcome before any argument gets weighed.

Inside the pipeline from voice to verdict

Diagram: Where Bias Enters the Pipeline: Voice to Verdict. Visualizes: Show the two-stage pipeline from spoken argument to AI verdict: Stage 1 is the ASR layer (speech-to-text), Stage 2 is the language model (text scoring).

A spoken round judged by AI has to move through at least two separate systems before it gets a winner. The first is a speech-to-text engine, the automatic speech recognition (ASR) layer, which takes the raw audio of a debater's speech and turns it into a written transcript. The second is a language model that reads that transcript and scores it on logic, evidence, clarity, and rebuttal. These two systems do different jobs, and the second one works from text alone, never touching the audio. It reads text, and only text.

Because the judge's verdict is built entirely on what the transcription layer decided to write down, any transcription error becomes the record the judge works with. If the ASR system drops a word, mishears a phrase, or substitutes one term for another, that error does not get corrected later. The word is simply gone from the record the judge works with, as if it had never been said.

A newer architecture, the end-to-end audio-language model, collapses these two steps into one system that processes raw speech directly without producing an explicit transcript. So you remove one point of failure, but you add a bias profile of its own, separate from what pipeline ASR produces, and this piece takes that up later.

Text-based debate platforms avoid this entire layer. When debaters type their arguments rather than speak them, there is no audio to transcribe and no ASR system to introduce error. The fairness case for AI judging is strongest in that setting, because the step where bias enters a spoken round does not exist.

Spoken rounds carry no such shortcut. Every word must survive the trip from voice to text before the judge can score it.

Why ASR is the layer where score gaps originate

Accent and dialect bias in spoken AI debate judging is a transcription problem, not a judging problem. It enters the pipeline at the ASR layer, before any argument reaches the model that scores logic or rebuttal. ASR systems are trained predominantly on Standard American English, and they produce measurably higher error rates when the speaker uses a different variety of English, whether that variety is regional, ethnic, or shaped by a first language other than English.

The imbalance behind that gap is built into how these systems were made. Commercial ASR models have been trained and tested on speech corpora that disproportionately represent white, younger, non-disabled, native English speakers, so the systems learned one acoustic profile far better than others. Teleki and colleagues' cross-community speech AI research agenda frames this as a foundational flaw: speech AI works from an incomplete model of communication, an incomplete model of identity, and metrics that measure the wrong things. Speech carries social identity in a way text never does. Accent, intonation, and dialect are inseparable from the words themselves, so a system that transcribes speech is always also making a judgment, whether anyone designed it to or not, about who is speaking.

A court reporter and a judge fail in different ways, and that difference is a useful comparison here. A court reporter who mishears testimony produces a transcript that does not match what was said, and whatever errors land in that transcript become the official record, regardless of how clearly the witness actually spoke. A judge who later misapplies the law to that record is making a mistake of legal reasoning. AI debate judging has the same two failure points, but only one of them is where the trouble actually starts. The scoring model applies its rubric correctly to whatever text arrives. The text that arrives is where accent and dialect bias does its damage.

The downstream effect is concrete. A debater whose speech gets misrecognized arrives at the scoring stage with an already damaged argument record: dropped words, mangled grammar, rebuttals that never made it into text. The judge then scores that damage as if it were the debater's own argument, because the judge cannot tell a weak point from a point the transcription layer simply lost.

Which speakers are most affected across systems

The pattern is not random noise scattered across individual platforms. Major commercial ASR systems produce this pattern consistently, so a speaker disadvantaged by one system is very likely to be disadvantaged by the rest of them too.

The most extensively documented gap affects speakers of African American English. Koenecke and colleagues, in widely cited research, found substantially higher error rates for Black speakers than for white speakers across Amazon, Apple, Google, IBM, and Microsoft's ASR systems, a result that has become the field's standard reference point. That finding is not isolated. A sociophonetic analysis of the Pacific Northwest English Corpus found a 33% relative gap in word error rate between African American and Caucasian American speakers, and mixed-effects models backed this up at a high level of statistical significance (p < 0.001). Two separate datasets and two separate research teams found the same direction of disparity, making this a systematic feature of how these models were built.

Non-native English speakers face a related but separately rooted disadvantage: it depends on how far their first language sits from English structurally. Cheng, Clemmensen, and Das, working at the Technical University of Denmark, found that greater linguistic distance between a speaker's first language and English predicts higher ASR error rates, an association that held across multiple datasets and model architectures at a high level of statistical significance (p < 0.001). Their analysis of the models' internal representations found something sharper than a data gap: the latent space of most evaluated ASR architectures shows speakers segregating by first-language background at deeper acoustic layers. The models fail to normalize accented speech because they internally sort speakers by linguistic background as part of how they process audio, so the bias lives in the architecture, not only in what the training data contained. A separate analysis of Whisper and Seamless-M4T found large swings in word error rate across 26 different accent groups, so this is not a problem confined to one model family.

Regional dialects add a third dimension, apart from ethnicity or first language. Serditova, Tang, and Steffens studied Newcastle English and found that ASR errors track specific phonological, lexical, and morphosyntactic features of the dialect far more closely than they track any social factor about the speaker. Two of the systems tested in parts of that study, Google's and Deepgram's, produced error rates high enough that the researchers had to exclude them from certain parts of the analysis.

Taken together, these three strands describe one structural handicap, not three unrelated findings. A debater who is not young, white, a native English speaker, and a speaker of Standard American English faces measurably worse transcription at the exact layer that determines what the judge gets to read, and these disadvantages compound rather than cancel out when a speaker falls into more than one of these categories at once.

Diagram: Consistent Error Gaps Across Speaker Groups and ASR Systems. Visualizes: Show three documented word-error-rate disparities as magnitude callouts or a ranked gap chart: (1) African American vs.

What corrupted transcripts do to scored arguments

A transcription error in a debate round does not just make the same argument a little noisier. It corrupts the specific structures an AI judge is built to evaluate: the rebuttal, the evidence chain, the logical sequence connecting one claim to the next.

AI judges score dimensions like logical coherence, response relevance, and rebuttal strength, and every one of those dimensions depends on an accurate, complete record of what each speaker said and in what order. A rebuttal that gets dropped from the transcript reads to the judge as a conceded point, even if the debater delivered it clearly. A conditional claim that loses its second half in transcription, "if X then Y" becoming just "if X", removes the exact inferential link the judge would otherwise credit. If a piece of evidence gets garbled, the claim it supports looks unsupported on the page, even though it was fully supported in speech.

Kadoma, Shrivastava, and Naaman, working across Cornell University, Carnegie Mellon University, and Cornell Tech, ran an experiment in 2026 showing that ASR subtitle errors consistently lower both speaker evaluations and content evaluations, for every speaker tested. Because speakers with non-standard accents already receive worse transcription at a higher baseline rate, this effect lands on them twice: once from the transcription error itself, and again from the lower evaluation that error produces, affecting judgments of both the argument and the person making it.

This exposes a premise that AI debate judging depends on without quite stating it: that the weaker argument loses. The judge's rubric applies its scoring criteria faithfully to the text in front of it without malfunctioning. The text in front of it just is not a reliable record of what the better argument was.

The human cost that error metrics alone don't capture

Error rate statistics describe what happens to a transcript. They do not describe what a speaker has to do, in real time, to keep a flawed system from working against them.

Speakers who face ASR bias often adapt their own speech to compensate for the system's limitations, and that adaptation costs them time in a round where a debater speaking into a system built around their own accent pays nothing. Liang and Beckford Wassink, presenting at FAccT 2026 out of the University of Washington, studied four U.S. dialect communities and found that most participants feel these technologies fail to account for their cultural backgrounds and require constant, active adjustment just to get basic functionality out of them. African American and non-native English speakers in that research reported real dissatisfaction and mistrust toward ASR technology, pointing to frequent misunderstandings and interruptions in how the system handled their speech.

In a debate round, the most costly of these adaptations is hyper-articulation: slowing down, deliberately modifying pronunciation, or switching registers mid-speech to try to get the system to recognize what's being said. Every one of those adjustments changes the pace and delivery of the argument itself, and pace and delivery are what AI judges score under persuasion and rhetoric. If a debater has to spend effort just being heard correctly, less of it is left to make the argument land. That tradeoff does not appear in any error rate; it appears in the gap between the argument a debater is capable of making and the argument they actually deliver once part of their attention is occupied by managing a system not built with their voice in mind.

Why a pipeline system doesn't fully solve the problem

There's a genuine technical argument that explicit transcription, the pipeline approach, reduces some of this bias. The 2026 BiasInEar benchmark, produced by Wei, Liao, Chang, Huang, and Chen for EACL Findings, found that pipeline systems (transcription-first) show higher agreement and lower sensitivity across language, accent, and gender compared to end-to-end systems. So converting speech to text seems to suppress some of the speaker-dependent variability that would otherwise reach the scoring stage.

That finding describes what happens once a transcript already exists, not what happens while it's being produced. A pipeline system can be as consistent as it likes in how it scores a transcript, and none of that consistency repairs a transcript that misrecognized a quarter of what a Newcastle English speaker said, or dropped a rebuttal delivered in African American English. The bias in a pipeline system lives in the transcript. So making the judge behave more consistently once it reads that transcript does nothing to fix what the transcript contains.

End-to-end models skip the transcript step, but they carry their own bias profile. A 2026 study of SpeechLLM voice cloning found directional disparities tied to accent, with Eastern European-accented speech receiving lower helpfulness scores, an effect most pronounced for female-presenting voices. The same research found that human evaluators detected these disparities more readily than the LLM judges did, which suggests end-to-end systems can mask biases that a human listener would catch immediately. Choosing between pipeline and end-to-end architecture is a real design decision with real tradeoffs on both sides, but neither one addresses the actual source of the problem: ASR and end-to-end models alike are trained on data that under-represents the speech varieties producing these gaps.

What platforms can do: diagnosis, data, and transparency

The root cause here, under-representation of non-mainstream speech varieties in training data, is a solvable problem, and the research points to specific interventions that operate at the system level rather than asking speakers to keep adjusting to the system.

Manual, sociolinguistic error analysis is the necessary first step. Serditova, Tang, and Steffens' Newcastle English work shows that systematic review of transcription errors can pinpoint the exact phonological, lexical, and morphosyntactic features driving misrecognition, and that this kind of analysis can feed directly into targeted data augmentation and the design of more inclusive acoustic models. Their finding that regional pronouns like "yous" and "wor" get systematically misrecognized is a useful illustration of how narrow and fixable these errors can be once someone has actually gone looking for them. So a platform running AI-judged spoken rounds has every reason to run this kind of diagnostic against its own transcription layer instead of waiting for an external benchmark to surface a gap.

Debaters evaluating whether a spoken-round AI judge can be trusted can ask what ASR system sits underneath the scoring layer, what accent and dialect groups it has been tested against, and whether the platform publishes any of that error-rate data. A platform that cannot answer those questions is asking debaters to trust a transcript they have no way to check. Text-based judging removes the transcription layer entirely, so it sidesteps the whole question, which is why platforms like ArguFight, scoring typed rounds with published written reasoning, make a different and more verifiable kind of fairness claim than any spoken-round system can currently make. For spoken debate to reach that same standard of trust, the work has to happen at the transcription layer first, because that is where the gap between what was said and what gets judged actually opens up.

Sources

  1. "This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias
  2. Linguistic Distance Segregates Latent Representations in Automatic Speech Recognition Systems
  3. Lost in Transcription: Subtitle Errors in Automatic Speech Recognition Reduce Speaker and Content Evaluations
  4. Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
  5. A Cross Community Agenda for Speech AI
  6. A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English

More in Bias Auditing and Fairness