Patientdesk Labs · Release 2026.2

AI Voice Arena

Listeners hear two AI voices read the same sentence and pick one, never seeing who made either. Two things make this board different from the usual arena: who votes, and what they were asked.

Read the full report → English and Turkish today. More languages next round — tell us which one you need.
12,245Blind comparisons
56Voices tested
16Companies
224Listeners
25 Jul 2026Results frozen
01 · Who votes A vetted panel, not the open web 224 paid native speakers, every one cleared through onboarding before they could take listening work. There is no public vote button on this page. 02 · Better at what? Four questions, not “which is better” Fish Audio’s s2-pro is 5th of 44 on sounds human and 40th of 44 on clear and correct. Pick your question below.
Language · more soon
Question
Group by
# Model Voice 95% range People
95% range score bottom of leader’s range clearly behind the leader average voice (1000)

01 Who votes

A vetted panel, not the open web

Not whoever shows up. All 224 listeners are paid contributors from joinvoicedata.com who had already cleared onboarding — an approved recording, or an approved application — before any listening work was offered to them. Each one judged only their native language. There is no vote button on this page, so no ranking here moves because a link went around. The usual format is open by design: Hugging Face’s TTS Arena, for one, takes a vote from any account thirty days old.

Native speaker only Onboarding cleared Paid per comparison No public sign-up
9.3% of the hidden same-clip controls still came back with a confident winner picked — 149 of 1,597. Vetting is not a guarantee, and a vetted panel is still not a representative cross-section of either language.
02 Better at what?

Four questions, not “which is better”

“Better” never says better at what, so a one-question board hands you a single number with the trade-off already averaged out of it. We ask four separate questions and never mix them inside a sitting: overall preference, sounds human, clear and correct, rhythm and expression. In aggregate the four broadly agree. Per voice they do not, and per voice is what you ship — 33 of the 44 English voices move ten places or more depending on which one you read.

Fish Audio s2-pro · one voice, two questions
1st44th
Sounds human #5
Clear and correct #40
Same clips, same listeners, same week. Switch the Question filter above to move between these boards.

Scores come from a standard head-to-head method called Bradley–Terry, centred so the average voice in each language sits at 1000. A listener answers one of the four questions for a whole sitting of 25 comparisons and the questions are never mixed inside a sitting, so the 95% ranges are worked out by resampling listeners rather than individual clicks, because one person voting 25 times is not 25 independent opinions. Results are frozen once voting closes; we do not publish a moving scoreboard while people are still voting.

Read the full report and method →  ·  Download the raw results →  ·  Get paid to listen →