People listened to pairs of AI voices reading the same sentence and picked one, never seeing who made either. This is round two: 12,245 comparisons, 4.6 times the size of our pilot, from a panel we vetted before they voted and split across four questions instead of one.
s2-pro is 5th of 44 on sounding human and 40th of 44 on getting the words right. One number would have averaged that away.
What we did differently, what changed since the pilot, how the voting works, what round two found, and what we still cannot tell you.
Summary
Most AI voice companies show you a demo page. Very few tell you how they tested, how confident they are, or how the voice holds up in a language other than English. We needed to pick a voice for a phone agent that talks to real patients, could not find that information anywhere, so we measured it ourselves.
Round 2026.2 covers 44 English voices and 28 Turkish voices from 16 companies. Two voices read the same sentence, a listener picks one, and nobody sees a brand name until voting closes. Every score ships with an error bar, and the full results file is a public download.
The format is not ours — blind pairwise voting has been the standard for years. Two things inside it are. Every vote came from a vetted panel of paid native speakers rather than from anyone who could reach the page, and we asked four separate questions rather than one. Both changed what came out, and the rest of this report is mostly the consequences.
What We Did Differently
Blind pairwise voting is a solved format and we did not reinvent it. Hugging Face's TTS Arena has run it for years: hide the brand, play two clips, count the wins, publish an Elo table. We kept all of that. We changed who votes and what they are asked.
Neither change is free. A closed paid panel fills slowly and stays small: 12,245 votes, not millions. Four questions split those votes four ways, so every error bar is wider than it would be if we pooled everything into one ranking. And a vetted panel is still not a representative sample of either language — people who work on AI training data listen more carefully than a patient on a phone call does.
We took both costs because the alternative is a number nobody can interpret: an unknown population answering an undefined question, updating live. We would rather publish a smaller frozen table and tell you exactly who produced it.
What Changed Since The Pilot
Our July pilot was honest about being too small: 2,675 votes across 40 English voices meant 21 of the 22 ranked entries overlapped the leader, and we said plainly that we could not rank English. Round two is 4.6 times bigger, and the picture is different in three specific ways.
| Pilot 2026.1 | Round 2026.2 | |
|---|---|---|
| Votes counted | 2,675 | 12,245 |
| Listeners | 44 | 224 |
| Voices | 66 from 13 companies | 72 from 16 companies |
| Voices with enough votes to rank | 48 of 66 | all 72 |
| Typical English error bar | 340 points | 140 points |
| English voices clearly behind the leader | 1 of 21 | 20 of 43 |
First, English separates now — though not at the top. The error bars roughly halved, and 20 of the 44 voices are now measurably behind the leading pack instead of statistically tangled with it. The other 24 still overlap each other.
Second, the cross-language finding survived and got stronger. It was the one pilot result we said we believed, and more data has not softened it.
Third, and least comfortably, one pilot finding did not survive. We reported that the four listening questions measured genuinely different things, pointing at a near-zero correlation between “sounds human” and overall preference in Turkish. With four times the data that correlation is 0.89. It was noise. We were measuring our own sample size.
| How closely each question tracks overall preference | English pilot | English now | Turkish pilot | Turkish now |
|---|---|---|---|---|
| Sounds human | 0.46 | 0.61 | 0.08 | 0.89 |
| Clear and correct | 0.52 | 0.76 | 0.68 | 0.95 |
| Rhythm and expression | 0.41 | 0.76 | 0.54 | 0.87 |
Splitting the questions is still the right call, and not only because it is how you find out they agree. Agreement in aggregate is not agreement per voice: the same English data has 33 of 44 voices moving ten or more places depending on the question, and the two questions that track each other least — “sounds human” and “clear and correct”, at 0.59 — are exactly the two a buyer has to trade off. Better At What goes through it. What we no longer claim is that the questions measure unrelated things, and anyone who quoted our pilot on that should stop.
How It Works
The point of the setup is to test voices, not brand loyalty or listener patience. Nothing about the method changed between rounds, which is what makes the two comparable.
Who does the listening
Listeners are paid contributors on joinvoicedata.com, and they only reach listening work after clearing the platform's onboarding — an approved voice recording, or an approved application. Most have done linguistics or AI training-data work before. Nobody judges a language they did not declare as native, and there is no open sign-up: the arena page has no vote button, so the panel is the panel. Round two was 224 people across 556 completed sittings, paid per comparison.
Both voices read the same sentence
Every comparison uses identical text, from a fixed bank of 24 sentences — 12 English, 12 Turkish — covering plain statements, questions, lists, names, numbers and dates, abbreviations, long winding sentences, and emotional lines.
Nobody sees the brand names
The listener's browser gets an audio file and its length, and nothing else. File names are random IDs, so no company or model name appears in the page, the network traffic, or anywhere a curious listener could dig. Buttons say “First voice” and “Second voice”. The map from clip to company lives in a separate file that is only opened after voting closes.
Four questions, asked one at a time
A listener gets 25 comparisons in a sitting, and all 25 ask the same one of four questions: which voice you prefer overall, which sounds more like a person, which is clearer and gets the words right, or which has the better rhythm and expression. They are never mixed inside a sitting, because an impression formed while judging rhythm should not decide the next answer about pronunciation. The server rejects any vote where the listener did not play at least 60% of both clips, so a modified browser cannot fake it. The two players cannot run at once. Replays are unlimited. There is no skip — you pick a voice, say they sound the same, or report an audio problem. Nobody sees the same pair twice, and nobody votes more than once on a given question in a given language.
Hidden checks
Three of the 25 comparisons are traps: the same clip played on both sides, at positions 5, 13 and 21. They look completely normal, both sides getting their own download link for the identical file. The honest answers are “about the same” or “audio problem”; picking a winner counts as a miss. Traps are paid like any other comparison, which is what keeps them hidden, and they never count toward anyone's score.
How the scores work
We use a standard head-to-head method called Bradley–Terry, which turns a pile of pairwise results into a strength number per voice. A tie counts as half a win. Scores are centred so the average voice in each language sits at 1000. Error bars come from resampling listeners rather than individual clicks, because one person voting 25 times is not 25 independent opinions. A voice needs 20 votes to earn a rank; in this round every voice cleared that easily.
What We Collected
Voting ran from 20 to 25 July 2026 and closed at 556 completed sittings.
English Results
This is what the pilot could not give us, with one caveat. The field is 44 voices, the bottom 20 are now measurably behind, and the top 24 are still a crowd.
Google Gemini's gemini-3.1-flash-tts-preview leads at 1138, but its range overlaps every voice down to rank 24. The honest reading is that the English field splits in two: a broad leading pack of 24 voices we cannot order internally, and 20 that are clearly behind it. That is a real gain over the pilot, where only one voice was separable from the leader — but it is separation of the bottom half, not a winner at the top.
| Rank | Company | Model | Voice | Score | 95% range | Votes |
|---|---|---|---|---|---|---|
| 1 | Google Gemini | gemini-3.1-flash-tts-preview | Puck | 1137.8 | 1075–1205 | 93 |
| 2 | Smallest | lightning_v3.1_pro | nolan | 1125.4 | 1055–1209 | 86 |
| 3 | Cartesia | sonic-3.5-2026-05-04 | db6b0ed5 | 1123.7 | 1069–1184 | 100 |
| 4 | Speechify | simba-3.2 | geffen_32 | 1121.2 | 1051–1191 | 107 |
| 5 | xAI | grok-voice-tts | altair | 1107.3 | 1036–1178 | 98 |
| 6 | Speechify | simba-3.2 | dominic_32 | 1099.8 | 1036–1161 | 98 |
| 7 | Resemble AI | chatterbox | 6e37aa15 | 1095.1 | 1028–1175 | 97 |
| 8 | Google Gemini | gemini-3.1-flash-tts-preview | Kore | 1092.4 | 1033–1161 | 116 |
| 9 | Cartesia | sonic-3.5-2026-05-04 | 630ed21c | 1083.2 | 1015–1148 | 89 |
| 10 | Gradium | default | YTpq7expH9539ERJ | 1077.0 | 999–1153 | 112 |
| 11 | MiniMax | speech-2.8-hd | English_Trustworth_Man | 1062.7 | 1005–1127 | 88 |
| 12 | Smallest | lightning_v3.1_pro | chelsea | 1061.4 | 1007–1117 | 109 |
| 13 | Cartesia | sonic-3 | db6b0ed5 | 1060.5 | 997–1138 | 111 |
| 14 | Resemble AI | chatterbox | 819fcc57 | 1060.4 | 1003–1121 | 108 |
| 15 | ElevenLabs | eleven_v3 | iP95p4xoKVk53GoZ742B | 1058.7 | 988–1140 | 91 |
| 16 | StepFun | step-tts-2 | lively-girl | 1042.6 | 971–1108 | 98 |
| 17 | Hume AI | 1 | Female Meditation Guide | 1037.5 | 981–1102 | 104 |
| 18 | Gradium | default | LFZvm12tW_z0xfGo | 1029.9 | 966–1088 | 95 |
| 19 | xAI | grok-voice-tts | ara | 1028.8 | 974–1092 | 108 |
| 20 | Async | async_flash_v1.0 | cca0e076 | 1023.3 | 958–1091 | 106 |
| 21 | Google Gemini | gemini-2.5-pro-preview-tts | Kore | 1022.4 | 973–1079 | 111 |
| 22 | MiniMax | speech-2.8-hd | English_CalmWoman | 1021.2 | 951–1091 | 110 |
| 23 | Async | async_flash_v1.5 | e5a67eaf | 1018.4 | 942–1092 | 83 |
| 24 | Async | async_flash_v1.0 | 317bf805 | 1007.3 | 931–1076 | 90 |
| 25 | OpenAI | gpt-4o-mini-tts | nova | 1005.2 | 936–1072 | 101 |
| 26 | MiniMax | speech-2.8-turbo | English_CalmWoman | 1001.3 | 932–1067 | 114 |
| 27 | Hume AI | 2 | Female Meditation Guide | 992.7 | 931–1061 | 106 |
| 28 | Cartesia | sonic-3 | 630ed21c | 974.5 | 891–1048 | 93 |
| 29 | Rime | coda | masonry | 969.1 | 894–1036 | 77 |
| 30 | Fish Audio | s2-pro | e3cd3841 | 956.8 | 885–1031 | 108 |
| 31 | Async | async_flash_v1.5 | 0aef6559 | 956.2 | 885–1017 | 114 |
| 32 | Smallest | lightning_v3.1 | liam | 955.9 | 882–1031 | 97 |
| 33 | Smallest | lightning_v3.1 | avery | 954.7 | 884–1021 | 97 |
| 34 | Fish Audio | s2.1-pro-free | c2623f0c | 949.3 | 884–1015 | 104 |
| 35 | OpenAI | gpt-4o-mini-tts | onyx | 940.4 | 860–1024 | 84 |
| 36 | Rime | coda | astra | 936.2 | 862–997 | 107 |
| 37 | ElevenLabs | eleven_v3 | Xb7hH8MSUJpSbSDYk0k2 | 915.7 | 842–988 | 101 |
| 38 | Google Gemini | gemini-2.5-pro-preview-tts | Puck | 899.6 | 825–977 | 92 |
| 39 | StepFun | step-tts-2 | magnetic-voiced-male | 891.2 | 814–966 | 92 |
| 40 | Inworld | inworld-tts-2 | Bianca | 883.8 | 815–938 | 109 |
| 41 | Inworld | inworld-tts-2 | Callum | 874.5 | 799–946 | 87 |
| 42 | Inworld | inworld-tts-1.5-max | Bianca | 865.5 | 793–936 | 109 |
| 43 | Inworld | inworld-tts-1.5-max | Callum | 863.1 | 775–943 | 84 |
| 44 | Fish Audio | s2.1-pro-free | c5f56a6c | 616.2 | 399–726 | 86 |
Turkish Results
Turkish has a smaller field and a much clearer top. 19 of 27 challengers are clearly behind the leader.
gemini-3.1-flash-tts-preview Puck leads at 1302 with a range of 1219–1417, clear of every rival.| Rank | Company | Model | Voice | Score | 95% range | Votes |
|---|---|---|---|---|---|---|
| 1 | Google Gemini | gemini-3.1-flash-tts-preview | Puck | 1301.5 | 1219–1417 | 58 |
| 2 | Google Gemini | gemini-3.1-flash-tts-preview | Kore | 1194.9 | 1094–1302 | 66 |
| 3 | Google Gemini | gemini-2.5-pro-preview-tts | Kore | 1184.7 | 1105–1286 | 64 |
| 4 | Google Gemini | gemini-2.5-pro-preview-tts | Puck | 1183.4 | 1103–1293 | 60 |
| 5 | xAI | grok-voice-tts | ara | 1157.4 | 1083–1248 | 61 |
| 6 | MiniMax | speech-2.8-hd | Turkish_CalmWoman | 1150.6 | 1072–1252 | 66 |
| 7 | ElevenLabs | eleven_v3 | iP95p4xoKVk53GoZ742B | 1146.6 | 1068–1232 | 65 |
| 8 | ElevenLabs | eleven_v3 | Xb7hH8MSUJpSbSDYk0k2 | 1132.5 | 1044–1245 | 64 |
| 9 | Resemble AI | chatterbox-multilingual | 64ad5770 | 1128.4 | 1048–1222 | 60 |
| 10 | MiniMax | speech-2.8-turbo | Turkish_CalmWoman | 1120.7 | 1047–1216 | 62 |
| 11 | Resemble AI | chatterbox-multilingual | 45b91687 | 1093.4 | 1021–1174 | 57 |
| 12 | Fish Audio | s2.1-pro-free | 67890dc7 | 1065.2 | 984–1148 | 55 |
| 13 | xAI | grok-voice-tts | altair | 1061.5 | 964–1142 | 55 |
| 14 | MiniMax | speech-2.8-hd | Turkish_Trustworthyman | 1048.1 | 980–1128 | 56 |
| 15 | Fish Audio | s2-pro | 818b100d | 1017.0 | 936–1119 | 60 |
| 16 | Async | async_flash_v1.0 | 317bf805 | 967.1 | 884–1051 | 61 |
| 17 | Speechify | simba-multilingual | ayse | 920.7 | 832–1008 | 56 |
| 18 | Cartesia | sonic-3.5-2026-05-04 | 630ed21c | 906.1 | 810–979 | 66 |
| 19 | Cartesia | sonic-3.5-2026-05-04 | db6b0ed5 | 900.0 | 763–992 | 53 |
| 20 | Speechify | simba-multilingual | arda | 893.2 | 795–988 | 60 |
| 21 | Cartesia | sonic-3 | db6b0ed5 | 876.7 | 779–968 | 64 |
| 22 | Speechify | simba-multilingual | baris | 876.3 | 759–967 | 55 |
| 23 | OpenAI | gpt-4o-mini-tts | onyx | 826.5 | 719–928 | 63 |
| 24 | OpenAI | gpt-4o-mini-tts | nova | 825.4 | 729–902 | 59 |
| 25 | Async | async_flash_v1.0 | cca0e076 | 808.0 | 694–902 | 60 |
| 26 | Speechify | simba-multilingual | aylin | 763.0 | 623–866 | 66 |
| 27 | Cartesia | sonic-3 | 630ed21c | 730.1 | 619–815 | 61 |
| 28 | Fish Audio | s2.1-pro-free | fe96529b | 721.0 | 578–817 | 69 |
The Language Gap
Sixteen of the voices were put in front of listeners in both languages — identical company, model and voice ID. That is the cleanest possible test of whether quality travels, because nothing changes except the language being read.
It does not travel.
| Voice (identical in both languages) | English | Turkish | Direction |
|---|---|---|---|
Cartesia db6b0ed5 | #3 of 44 | #19 of 28 | falls in Turkish |
Cartesia db6b0ed5 | #13 of 44 | #21 of 28 | falls in Turkish |
Async cca0e076 | #20 of 44 | #25 of 28 | falls in Turkish |
Cartesia 630ed21c | #9 of 44 | #18 of 28 | falls in Turkish |
xAI altair | #5 of 44 | #13 of 28 | falls in Turkish |
Cartesia 630ed21c | #28 of 44 | #27 of 28 | falls in Turkish |
OpenAI nova | #25 of 44 | #24 of 28 | falls in Turkish |
Async 317bf805 | #24 of 44 | #16 of 28 | falls in Turkish |
OpenAI onyx | #35 of 44 | #23 of 28 | falls in Turkish |
Google Gemini Puck | #1 of 44 | #1 of 28 | falls in Turkish |
ElevenLabs iP95p4xoKVk53GoZ742B | #15 of 44 | #7 of 28 | rises in Turkish |
Google Gemini Kore | #8 of 44 | #2 of 28 | rises in Turkish |
xAI ara | #19 of 44 | #5 of 28 | rises in Turkish |
Google Gemini Kore | #21 of 44 | #3 of 28 | rises in Turkish |
ElevenLabs Xb7hH8MSUJpSbSDYk0k2 | #37 of 44 | #8 of 28 | rises in Turkish |
Google Gemini Puck | #38 of 44 | #4 of 28 | rises in Turkish |
One important exception, stated plainly: Google Gemini's gemini-3.1-flash-tts-preview Puck is ranked 1st in both languages. A voice can be good in two languages at once. The point is that most are not, and you cannot tell which from an English score.
Pooling all four questions per company gives each one enough comparisons for the cross-language picture to resolve sharply. This is the most useful thing in the report.
The two highlighted companies move in opposite directions, and neither move is small.
| Company | English | Turkish | What happened |
|---|---|---|---|
| Speechify | 63.6% — 1st of 16 | 37.3% — 8th of 10 | top to near-bottom |
| ElevenLabs | 45.4% — 12th of 16 | 64.4% — 3rd of 10 | near-bottom to top |
| Cartesia | 57.7% — 3rd of 16 | 35.5% — 9th of 10 | top to bottom |
| Google Gemini | 57.2% — 5th of 16 | 72.7% — 1st of 10 | strong to dominant |
| OpenAI | 46.0% — 11th of 16 | 30.3% — 10th of 10 | weak in both |
| xAI | 60.1% — 2nd of 16 | 65.3% — 2nd of 10 | the only steady one |
Cartesia repeats its pilot result almost exactly, which is the best evidence we have that the effect is real rather than an artefact of one round. Speechify is the new and starker case: first in English, eighth of ten in Turkish.
Six of the sixteen companies fielded nothing in Turkish at all — Smallest, Gradium, StepFun, Hume AI, Rime and Inworld. Turkish speakers choose from 28 voices against 44, and four of the top ten Turkish slots belong to Google.
| # | Company | Win rate | 95% range | Comparisons | Voices |
|---|---|---|---|---|---|
| 1 | Speechify | 63.6% | 60.2–66.9% | 799 | 2 |
| 2 | xAI | 60.1% | 56.8–63.5% | 821 | 2 |
| 3 | Cartesia | 57.7% | 55.3–60.2% | 1557 | 4 |
| 4 | Smallest | 57.4% | 55.0–59.8% | 1576 | 4 |
| 5 | Google Gemini | 57.2% | 54.8–59.6% | 1590 | 4 |
| 6 | Resemble AI | 54.3% | 50.8–57.8% | 779 | 2 |
| 7 | MiniMax | 52.6% | 49.9–55.4% | 1229 | 3 |
| 8 | Gradium | 50.6% | 47.2–54.0% | 818 | 2 |
| 9 | StepFun | 49.9% | 46.3–53.5% | 744 | 2 |
| 10 | Hume AI | 49.3% | 46.0–52.6% | 863 | 2 |
| 11 | OpenAI | 46.0% | 42.5–49.6% | 772 | 2 |
| 12 | ElevenLabs | 45.4% | 41.9–48.9% | 791 | 2 |
| 13 | Async | 43.5% | 41.0–45.9% | 1570 | 4 |
| 14 | Rime | 42.5% | 39.1–45.9% | 794 | 2 |
| 15 | Fish Audio | 37.3% | 34.6–40.1% | 1213 | 3 |
| 16 | Inworld | 36.1% | 33.7–38.4% | 1616 | 4 |
| # | Company | Win rate | 95% range | Comparisons | Voices |
|---|---|---|---|---|---|
| 1 | Google Gemini | 72.7% | 69.9–75.4% | 1022 | 4 |
| 2 | xAI | 65.3% | 61.0–69.5% | 485 | 2 |
| 3 | ElevenLabs | 64.4% | 60.3–68.5% | 520 | 2 |
| 4 | Resemble AI | 61.2% | 56.9–65.6% | 485 | 2 |
| 5 | MiniMax | 57.5% | 53.9–61.0% | 744 | 3 |
| 6 | Async | 40.2% | 35.9–44.4% | 509 | 2 |
| 7 | Fish Audio | 39.4% | 35.9–42.9% | 745 | 3 |
| 8 | Speechify | 37.3% | 34.3–40.4% | 972 | 4 |
| 9 | Cartesia | 35.5% | 32.6–38.5% | 999 | 4 |
| 10 | OpenAI | 30.3% | 26.2–34.4% | 477 | 2 |
Better At What?
This is the section that only exists because we asked four questions. Every number below comes from the same 12,245 votes; the only thing that changes is which question the listener was answering.
Start with the strongest case. In English, Fish Audio's s2-pro is 5th of 44 on sounding human and 40th of 44 on getting the words right. That is not a wobble inside the error bars: on “sounds human” its range overlaps the leader's, and on “clear and correct” its entire range (819–963) sits below the bottom of the leader's range at 1090. The same clips, the same listeners, the same week. It sounds like a person and it mispronounces things.
| English voice | Overall preference | Sounds human | Clear and correct | Rhythm and expression |
|---|---|---|---|---|
Fish Audio s2-pro | #30 | #5 | #40 | #28 |
Cartesia sonic-3 630ed21c | #28 | #2 | #30 | #15 |
Smallest lightning_v3.1 avery | #33 | #4 | #14 | #11 |
Gradium YTpq7expH9539ERJ | #10 | #34 | #31 | #25 |
Hume AI Female Meditation Guide | #27 | #31 | #10 | #21 |
xAI grok-voice-tts altair | #5 | #1 | #7 | #4 |
Google Gemini gemini-3.1-flash-tts-preview Puck | #1 | #11 | #11 | #2 |
Across the whole English field, 33 of 44 voices move ten or more places between their best and worst question. The median voice moves 13 places; the widest moves 35. Ask one question and you get one of those columns, with no way to know which.
The two questions that disagree most are the two that matter most commercially. “Sounds human” and “clear and correct” track each other at 0.59 in English, against 0.81 for clarity versus rhythm. A demo reel is won on the first; a phone call that reads back an address is won on the second.
xAI owns one question, Google owns the others
The clearest company-level split in the data. xAI's altair is the outright top of the English “sounds human” board at 1155, with 26 of the 44 voices clearly behind it — and in Turkish, xAI's ara tops that same question at 1188. Whatever xAI is doing about sounding like a person, it travels across both languages we tested.
It does not win the rest. In Turkish, Google Gemini leads three of the four questions outright, and the single question it does not lead is the one xAI takes.
| Turkish, by company | Overall preference | Sounds human | Clear and correct | Rhythm and expression |
|---|---|---|---|---|
| Google Gemini | 75.0% — 1st | 64.5% — 2nd | 75.8% — 1st | 75.8% — 1st |
| xAI | 65.1% — 3rd | 68.0% — 1st | 69.1% — 2nd | 58.7% — 4th |
Being careful about that Turkish naturalness row: 68.0% and 64.5% have overlapping ranges, so read it as xAI edging ahead rather than beating Google. The voice-level result behind it is the harder evidence, and it points the same way in both languages.
The English winner makes the same point from the other side. Google Gemini's gemini-3.1-flash-tts-preview Puck is 1st overall and 2nd on rhythm, but only 11th of 44 on sounding human and 11th on clarity. It wins English on delivery, not on humanity. An arena that asked one question would have told you it is simply the best English voice, and that is not what the listeners said.
Two honest qualifications. First, in aggregate the four questions do broadly agree on the ordering — between 0.61 and 0.76 against overall preference in English — and we are not reviving the retired pilot claim that they measure unrelated things. The disagreement is per voice, which is unfortunately the level at which you have to choose one. Second, Turkish is much tighter than English: the median Turkish voice moves 6 places across the four questions and the widest moves 12, against 13 and 35 in English. With 28 voices and fewer votes, we would expect a smaller field to spread less; do not read the Turkish calm as a finding.
Voice Versus Model
People compare companies and models, but what you ship is one voice. Within a single model the spread between voices is still large enough to dominate the choice of model.
| Model | Best voice | Worst voice | Gap |
|---|---|---|---|
Fish Audio s2.1-pro-free (English) | #34 c2623f0c — 949 | #44 c5f56a6c — 616 | 333 points |
Fish Audio s2.1-pro-free (Turkish) | #12 67890dc7 — 1065 | #28 fe96529b — 721 | 344 points |
StepFun step-tts-2 (English) | #16 lively-girl — 1043 | #39 magnetic-voiced-male — 891 | 151 points |
ElevenLabs eleven_v3 (English) | #15 iP95p4xo — 1059 | #37 Xb7hH8MS — 916 | 143 points |
Fish Audio's English pairing is the extreme case: the same model, the same language, 333 points and ten ranks apart, with one of the two finishing last of all 44 voices. Test the voice you are going to ship, not the company that makes it.
Do Listeners Pay Attention
A listening test is only as good as the people doing the listening. Unlike the pilot, this round had the hidden traps running the whole time, so the answer is measured rather than assumed.
So roughly one time in ten, someone hears the exact same clip twice and still picks a favourite. That is the cost of asking humans to compare things, and it is why we put the traps in. Any listening test without checks like these is carrying about the same noise and does not know it.
Who We Excluded, And Why
Being transparent about this matters more than the result looking clean, so here is exactly what happened before we froze the numbers.
Two contributors tripped the automatic disqualification rule, which needs at least 6 controls seen, at least 3 misses, and a miss rate of 50% or higher. Their votes are out.
A further 51 sat on a softer hold, triggered by two misses. We looked at each one against the 9.3% population miss rate:
| Group | People | Miss rate | Chance of that record if listening honestly | Decision |
|---|---|---|---|---|
| Missed 2 of 2 controls | 23 | 100% | 0.9% | excluded |
| Missed 2 of 3 | 7 | 67% | 2.4% | excluded |
| Missed 2 of 4 | 1 | 50% | 4.6% | excluded |
| Missed 2 of 5 to 2 of 12 | 20 | 17–40% | 7–31% | readmitted |
The 20 whose records are statistically consistent with honest listening were readmitted, returning 1,103 votes to the ranking. The 31 whose records are not — 23 of whom missed every single control they were shown — stay excluded, holding back 555 votes.
That line is a judgement call and we are showing our work rather than hiding it. Note that the automatic rule alone would have cleared all 31, purely because they had not yet been shown 6 controls each. A rule that needs six trials cannot catch someone who fails their first two.
What This Cannot Tell You
We would rather list these than have a reader find them.
Check It Yourself
The results file is public and signed. It is the same file that draws the leaderboard, so there is no private version.
| What | Where |
|---|---|
| Browse the leaderboard | The AI Voice Arena — filter by language and question, group by voice or company |
| Full results file | joinvoicedata.com/api/public/tts-benchmark?download=1 |
| Paid listening work | joinvoicedata.com/tasks |
The file carries its own checksum and signature, when it was generated, and a plain description of the method. Each leaderboard reports how many votes it rests on and how many people were behind them.
Working with us
We run these blind English and Turkish tests on unreleased models and tell you where the pronunciation and rhythm problems are before your users find them. If you build AI voices and want to be in the next round, or want a private evaluation, get in touch through patientdesk.ai. If you think we got something wrong, tell us — we publish corrections next to the original rather than quietly editing it.