Askable Labs has launched VOICE-H, a benchmark for evaluating AI-generated speech against human recordings. The study found listeners preferred emotional fit over human likeness.
The benchmark compared nine text-to-speech models with original human recordings from the same conversations in nearly 4,500 blind comparisons. It used 100 clips selected from more than 40,000 real interviews on Askable's research platform. A total of 300 participants in Australia, New Zealand, the United Kingdom and the United States each completed 15 randomised head-to-head tests.
Google's Gemini 3.1 Flash Voice ranked first overall with an Elo score of 1101, ahead of the human recording on 1045 and Cartesia's Sonic 3.5 on 1041. Gemini also led the individual categories for naturalness, accuracy, and emotion and tone.
The results suggest a shift in how synthetic voices are judged. Rather than favouring voices that sounded most human, participants consistently preferred voices whose delivery matched the feeling of the words being spoken.
Human speech still ranked highly for naturalness, placing second behind Gemini with a score of 1062 to 1094. But the human recording fell to eighth for accuracy on 947, behind seven AI systems, suggesting listeners rewarded clear pronunciation and consistency even when a voice was less lifelike.
The widest gap appeared in emotion and tone. Gemini scored 1146 in that category, more than 110 Elo points ahead of the human recording on 1034.
Why voices lost
Participant feedback showed the most common complaint was that voices sounded robotic, drawing 590 negative mentions across almost every model tested. But listeners also marked down voices that sounded too theatrical or overly performative, indicating that both flatness and exaggeration reduced preference.
One model's pauses and filler words drew 52 complaints and no positive mentions, contributing to 41 lost votes. The study also found listeners often reacted badly when systems reproduced written disfluencies too literally, although they responded more favourably when hesitations appeared more naturally in the flow of speech.
Unlike many voice evaluations that rely on technical measures or simple preference rankings, VOICE-H collected written explanations for every choice. That created a qualitative dataset alongside the pairwise comparisons and allowed researchers to examine why listeners preferred one voice over another.
The approach also benchmarked systems against real human conversation rather than scripted prompts. Researchers used identity-verified panellists and published demographics, while rankings were calculated using an ordinal Bradley-Terry model weighted by the strength of each preference.
John Goleby, Chief Executive Officer of Askable Labs, said the findings challenge the common assumption that realism is the main target for voice systems.
"People aren't rewarding voices simply because they're realistic. They're rewarding voices that communicate the right emotion at the right moment," Goleby said.
"That's a fundamentally different challenge, and one that requires a different way of measuring quality."
Method and ranking
The benchmark covered leading text-to-speech systems from companies including Google, OpenAI and ElevenLabs. Each AI output was judged directly against the original human recording of the same speech, allowing the study to compare not only systems against one another but also their performance against a real conversational baseline.
The findings suggest naturalness alone did not determine listener preference. Participants were willing to overlook minor imperfections when a voice carried the right emotional tone, while highly natural voices lost favour when their delivery felt mismatched to the content.
Isaac Povey, who led the benchmark, said the main value of the system lay in the explanations behind the rankings rather than the rankings alone.
"Most benchmarks tell you which voice won. The more useful signal is why," Povey said.
"When you ask people to explain a preference, they aren't chasing realism for its own sake. They want a voice whose emotion fits the moment. That's what we built VOICE-H to measure."
The study offers a snapshot of a market in which AI voices are closing the gap with, and in some cases overtaking, human recordings on listener preference. In Askable Labs' test, the original human sample ranked behind Gemini overall and only narrowly ahead of the third-placed system.
For developers of voice assistants and speech tools, the data suggests the contest is moving beyond whether a synthetic voice can pass as human. In this benchmark, the stronger differentiator was whether a system could deliver speech with the emotional balance listeners expected from the situation.
Gemini's lead in emotion and tone, alongside the weak showing of the human recording in accuracy, underlined that point: listeners did not simply reward realism, but judged voices on whether they sounded clear, appropriate and convincing for the moment.