VoiceData Research · Public benchmark · Release 2026.2
Which AI voices sound best in English and Turkish?
A blinded, human-preference benchmark. Listeners compare two voices reading the exact same sentence — never seeing the provider, model, or voice name — and every score ships with its uncertainty.
- Languages
- English · Turkish
- Evaluation
- Blinded pairwise
- Ranking model
- Bradley–Terry
- Uncertainty
- Bootstrap 95% CI
Results
Human preference, question by question
Frozen July 25, 2026
English · Sounds human
“Which voice sounds more like a real person?”
- 2,255
- judgments
- 104
- listeners
- 466
- voice pairs
| # | Voice system | Score · 95% CI438–1277 | BT score | Judgments | Listeners |
|---|---|---|---|---|---|
| 1 | xAI grok-voice-tts · altair | 1155.2 1091.3–1225.9 | 100 72% wins | 61 | |
| 2 | Cartesia sonic-3 · 630ed21c-2c5c-41cf-9d82-10a7fd668370 | 1137.3 1074.9–1221.9 | 86 69% wins | 58 | |
| 3 | Speechify simba-3.2 · dominic_32 | 1115.7 1054.6–1184.8 | 98 65% wins | 67 | |
| 4 | Smallest lightning_v3.1 · avery | 1114.4 1045.9–1186.9 | 118 65% wins | 69 | |
| 5 | Fish Audio s2-pro · e3cd384158934cc9a01029cd7d278634 | 1110.6 1048.3–1181.3 | 110 65% wins | 67 | |
| 6 | StepFun step-tts-2 · lively-girl | 1093.5 1032.9–1167.6 | 108 63% wins | 72 | |
| 7 | ElevenLabs eleven_v3 · iP95p4xoKVk53GoZ742B | 1093.1 1027.9–1167.2 | 98 63% wins | 65 | |
| 8 | Smallest lightning_v3.1_pro · chelsea | 1085.4 1028.2–1146.6 | 120 62% wins | 72 | |
| 9 | Speechify simba-3.2 · geffen_32 | 1081.8 1021.5–1149.0 | 110 60% wins | 67 | |
| 10 | Google Gemini gemini-3.1-flash-tts-preview · Kore Latest | 1080.4 1030.3–1146.9 | 120 60% wins | 73 | |
| 11 | Google Gemini gemini-3.1-flash-tts-preview · Puck Latest | 1078.7 1018.7–1142.7 | 86 60% wins | 59 | |
| 12 | Google Gemini gemini-2.5-pro-preview-tts · Kore | 1074.3 1011.4–1132.0 | 115 60% wins | 73 | |
| 13 | MiniMax speech-2.8-hd · English_Trustworth_Man | 1069.7 1001.3–1154.0 | 95 59% wins | 64 | |
| 14 | Cartesia sonic-3 · db6b0ed5-d5d3-463d-ae85-518a07d3c2b4 | 1068.7 1010.6–1138.7 | 108 59% wins | 71 | |
| 15 | Smallest lightning_v3.1_pro · nolan | 1058.4 991.7–1125.2 | 84 59% wins | 55 | |
| 16 | Fish Audio s2.1-pro-free · c2623f0c075b4492ac367989aee1576f Latest | 1055.0 984.7–1132.4 | 110 56% wins | 67 | |
| 17 | xAI grok-voice-tts · ara | 1052.9 999.5–1118.0 | 110 56% wins | 62 | |
| 18 | Cartesia sonic-3.5-2026-05-04 · db6b0ed5-d5d3-463d-ae85-518a07d3c2b4 Latest | 1033.3 966.9–1099.7 | 105 55% wins | 69 | |
| 19 | Resemble AI chatterbox · 819fcc57 | 1027.2 971.5–1079.3 | 104 53% wins | 68 | |
| 20 | Cartesia sonic-3.5-2026-05-04 · 630ed21c-2c5c-41cf-9d82-10a7fd668370 Latest | 1020.8 962.4–1085.7 | 90 53% wins | 62 | |
| 21 | Smallest lightning_v3.1 · liam | 1018.7 942.8–1089.2 | 93 53% wins | 63 | |
| 22 | Resemble AI chatterbox · 6e37aa15 | 1011.4 936.4–1080.6 | 88 52% wins | 62 | |
| 23 | MiniMax speech-2.8-hd · English_CalmWoman | 1000.3 941.7–1061.1 | 117 49% wins | 69 | |
| 24 | Gradium default · LFZvm12tW_z0xfGo | 999.8 926.7–1070.8 | 85 51% wins | 59 | |
| 25 | Rime coda · astra | 988.2 921.3–1055.9 | 118 48% wins | 69 | |
| 26 | Hume AI 1 · Female Meditation Guide | 972.1 907.0–1034.4 | 108 45% wins | 68 | |
| 27 | Async async_flash_v1.5 · e5a67eaf-6e5a-4488-9fb9-4806bd7fea54 Latest | 971.3 892.4–1038.3 | 91 47% wins | 63 | |
| 28 | Async async_flash_v1.0 · cca0e076-94b9-4c6d-86b7-546168f11174 | 970.7 900.8–1033.5 | 112 46% wins | 71 | |
| 29 | OpenAI gpt-4o-mini-tts · onyx | 969.8 898.5–1038.7 | 88 45% wins | 58 | |
| 30 | Google Gemini gemini-2.5-pro-preview-tts · Puck | 967.3 892.4–1027.0 | 82 46% wins | 54 | |
| 31 | Hume AI 2 · Female Meditation Guide Latest | 965.3 894.5–1030.6 | 112 45% wins | 67 | |
| 32 | MiniMax speech-2.8-turbo · English_CalmWoman | 950.3 890.5–1019.6 | 112 42% wins | 67 | |
| 33 | Inworld inworld-tts-1.5-max · Bianca | 944.8 882.2–1003.4 | 117 41% wins | 67 | |
| 34 | Gradium default · YTpq7expH9539ERJ | 943.2 879.5–1002.7 | 121 41% wins | 72 | |
| 35 | Rime coda · masonry | 935.0 870.9–999.3 | 95 42% wins | 70 | |
| 36 | StepFun step-tts-2 · magnetic-voiced-male | 928.0 849.5–1003.1 | 87 41% wins | 59 | |
| 37 | OpenAI gpt-4o-mini-tts · nova | 917.3 856.4–976.4 | 106 37% wins | 66 | |
| 38 | Inworld inworld-tts-1.5-max · Callum | 902.4 811.3–985.4 | 79 37% wins | 58 | |
| 39 | Async async_flash_v1.0 · 317bf805-4b42-417b-9474-10807e2f67c9 | 902.3 815.3–991.9 | 86 37% wins | 58 | |
| 40 | Inworld inworld-tts-2 · Callum Latest | 901.1 826.7–962.7 | 97 37% wins | 68 | |
| 41 | Inworld inworld-tts-2 · Bianca Latest | 879.4 809.6–940.7 | 117 32% wins | 78 | |
| 42 | Async async_flash_v1.5 · 0aef6559-9098-4860-8224-13e038ab3aef Latest | 866.0 787.5–931.1 | 116 31% wins | 74 | |
| 43 | ElevenLabs eleven_v3 · Xb7hH8MSUJpSbSDYk0k2 | 861.5 779.8–926.3 | 114 30% wins | 74 | |
| 44 | Fish Audio s2.1-pro-free · c5f56a6cc2ec4fa8920cb4c5889a3fb7 Latest | 627.6 489.7–718.7 | 94 11% wins | 60 |
12,432 submitted → 12,245 ranked · 187 excluded (187 audio problems, 0 invalid or duplicate)
Scores are comparable only within one language and one question. Intervals are bootstrapped by contributor, so overlapping ranges mean the apparent order may not replicate.
Protocol
Designed to test voices, not brand recognition
The benchmark separates English from Turkish and reports each listening question independently. It does not compress unlike qualities into one global score, and no result is published before its release is frozen.
01Matched text
Both systems read the exact same Unicode sentence in every comparison. A frozen sentence bank covers statements, questions, numbers, names, and expressive phrasing.
02Blinded listening
Audio IDs, storage paths, order, and browser responses contain no provider, model, or voice name while votes are collected. Left–right order is randomized.
03One question at a time
A 25-comparison set keeps one listening focus: overall preference, human-like sound, clear pronunciation, or rhythm and expression.
04Language-qualified listeners
English comparisons go to English-speaking contributors and Turkish comparisons to Turkish-speaking contributors. Listener strata are retained for analysis.
05Tie-aware ranking
Pairwise judgments fit a Bradley–Terry model; ties count as half-wins. Scores are centered at 1,000 within each language and question.
06Uncertainty shown
The 95% ranges are bootstrapped by contributor, not by individual click, so 25 correlated judgments from one listener are not treated as 25 independent people.
Need a private voice evaluation?
VoiceData can run blinded English and Turkish evaluations for unreleased models, diagnose pronunciation and naturalness gaps, and recruit consented speakers for targeted improvement data.