VoiceData Research · Public benchmark · Release 2026.2

Which AI voices sound best in English and Turkish?

A blinded, human-preference benchmark. Listeners compare two voices reading the exact same sentence — never seeing the provider, model, or voice name — and every score ships with its uncertainty.

Languages
English · Turkish
Evaluation
Blinded pairwise
Ranking model
Bradley–Terry
Uncertainty
Bootstrap 95% CI

Results

Human preference, question by question

Frozen July 25, 2026

Full results (JSON)

English · Sounds human

Which voice sounds more like a real person?

2,255
judgments
104
listeners
466
voice pairs
#Voice systemScore · 95% CI4381277BT scoreJudgmentsListeners
1

xAI

grok-voice-tts · altair

1155.2

1091.31225.9

100

72% wins

61
2

Cartesia

sonic-3 · 630ed21c-2c5c-41cf-9d82-10a7fd668370

1137.3

1074.91221.9

86

69% wins

58
3

Speechify

simba-3.2 · dominic_32

1115.7

1054.61184.8

98

65% wins

67
4

Smallest

lightning_v3.1 · avery

1114.4

1045.91186.9

118

65% wins

69
5

Fish Audio

s2-pro · e3cd384158934cc9a01029cd7d278634

1110.6

1048.31181.3

110

65% wins

67
6

StepFun

step-tts-2 · lively-girl

1093.5

1032.91167.6

108

63% wins

72
7

ElevenLabs

eleven_v3 · iP95p4xoKVk53GoZ742B

1093.1

1027.91167.2

98

63% wins

65
8

Smallest

lightning_v3.1_pro · chelsea

1085.4

1028.21146.6

120

62% wins

72
9

Speechify

simba-3.2 · geffen_32

1081.8

1021.51149.0

110

60% wins

67
10

Google Gemini

gemini-3.1-flash-tts-preview · Kore
Latest

1080.4

1030.31146.9

120

60% wins

73
11

Google Gemini

gemini-3.1-flash-tts-preview · Puck
Latest

1078.7

1018.71142.7

86

60% wins

59
12

Google Gemini

gemini-2.5-pro-preview-tts · Kore

1074.3

1011.41132.0

115

60% wins

73
13

MiniMax

speech-2.8-hd · English_Trustworth_Man

1069.7

1001.31154.0

95

59% wins

64
14

Cartesia

sonic-3 · db6b0ed5-d5d3-463d-ae85-518a07d3c2b4

1068.7

1010.61138.7

108

59% wins

71
15

Smallest

lightning_v3.1_pro · nolan

1058.4

991.71125.2

84

59% wins

55
16

Fish Audio

s2.1-pro-free · c2623f0c075b4492ac367989aee1576f
Latest

1055.0

984.71132.4

110

56% wins

67
17

xAI

grok-voice-tts · ara

1052.9

999.51118.0

110

56% wins

62
18

Cartesia

sonic-3.5-2026-05-04 · db6b0ed5-d5d3-463d-ae85-518a07d3c2b4
Latest

1033.3

966.91099.7

105

55% wins

69
19

Resemble AI

chatterbox · 819fcc57

1027.2

971.51079.3

104

53% wins

68
20

Cartesia

sonic-3.5-2026-05-04 · 630ed21c-2c5c-41cf-9d82-10a7fd668370
Latest

1020.8

962.41085.7

90

53% wins

62
21

Smallest

lightning_v3.1 · liam

1018.7

942.81089.2

93

53% wins

63
22

Resemble AI

chatterbox · 6e37aa15

1011.4

936.41080.6

88

52% wins

62
23

MiniMax

speech-2.8-hd · English_CalmWoman

1000.3

941.71061.1

117

49% wins

69
24

Gradium

default · LFZvm12tW_z0xfGo

999.8

926.71070.8

85

51% wins

59
25

Rime

coda · astra

988.2

921.31055.9

118

48% wins

69
26

Hume AI

1 · Female Meditation Guide

972.1

907.01034.4

108

45% wins

68
27

Async

async_flash_v1.5 · e5a67eaf-6e5a-4488-9fb9-4806bd7fea54
Latest

971.3

892.41038.3

91

47% wins

63
28

Async

async_flash_v1.0 · cca0e076-94b9-4c6d-86b7-546168f11174

970.7

900.81033.5

112

46% wins

71
29

OpenAI

gpt-4o-mini-tts · onyx

969.8

898.51038.7

88

45% wins

58
30

Google Gemini

gemini-2.5-pro-preview-tts · Puck

967.3

892.41027.0

82

46% wins

54
31

Hume AI

2 · Female Meditation Guide
Latest

965.3

894.51030.6

112

45% wins

67
32

MiniMax

speech-2.8-turbo · English_CalmWoman

950.3

890.51019.6

112

42% wins

67
33

Inworld

inworld-tts-1.5-max · Bianca

944.8

882.21003.4

117

41% wins

67
34

Gradium

default · YTpq7expH9539ERJ

943.2

879.51002.7

121

41% wins

72
35

Rime

coda · masonry

935.0

870.9999.3

95

42% wins

70
36

StepFun

step-tts-2 · magnetic-voiced-male

928.0

849.51003.1

87

41% wins

59
37

OpenAI

gpt-4o-mini-tts · nova

917.3

856.4976.4

106

37% wins

66
38

Inworld

inworld-tts-1.5-max · Callum

902.4

811.3985.4

79

37% wins

58
39

Async

async_flash_v1.0 · 317bf805-4b42-417b-9474-10807e2f67c9

902.3

815.3991.9

86

37% wins

58
40

Inworld

inworld-tts-2 · Callum
Latest

901.1

826.7962.7

97

37% wins

68
41

Inworld

inworld-tts-2 · Bianca
Latest

879.4

809.6940.7

117

32% wins

78
42

Async

async_flash_v1.5 · 0aef6559-9098-4860-8224-13e038ab3aef
Latest

866.0

787.5931.1

116

31% wins

74
43

ElevenLabs

eleven_v3 · Xb7hH8MSUJpSbSDYk0k2

861.5

779.8926.3

114

30% wins

74
44

Fish Audio

s2.1-pro-free · c5f56a6cc2ec4fa8920cb4c5889a3fb7
Latest

627.6

489.7718.7

94

11% wins

60
95% interval overlaps the leader — order not statistically settledInterval separated from the leader1,000 = language averageThe provider’s newest model here — older generations are listed unmarked

12,432 submitted → 12,245 ranked · 187 excluded (187 audio problems, 0 invalid or duplicate)

Scores are comparable only within one language and one question. Intervals are bootstrapped by contributor, so overlapping ranges mean the apparent order may not replicate.

Protocol

Designed to test voices, not brand recognition

The benchmark separates English from Turkish and reports each listening question independently. It does not compress unlike qualities into one global score, and no result is published before its release is frozen.

  1. 01Matched text

    Both systems read the exact same Unicode sentence in every comparison. A frozen sentence bank covers statements, questions, numbers, names, and expressive phrasing.

  2. 02Blinded listening

    Audio IDs, storage paths, order, and browser responses contain no provider, model, or voice name while votes are collected. Left–right order is randomized.

  3. 03One question at a time

    A 25-comparison set keeps one listening focus: overall preference, human-like sound, clear pronunciation, or rhythm and expression.

  4. 04Language-qualified listeners

    English comparisons go to English-speaking contributors and Turkish comparisons to Turkish-speaking contributors. Listener strata are retained for analysis.

  5. 05Tie-aware ranking

    Pairwise judgments fit a Bradley–Terry model; ties count as half-wins. Scores are centered at 1,000 within each language and question.

  6. 06Uncertainty shown

    The 95% ranges are bootstrapped by contributor, not by individual click, so 25 correlated judgments from one listener are not treated as 25 independent people.

Need a private voice evaluation?

VoiceData can run blinded English and Turkish evaluations for unreleased models, diagnose pronunciation and naturalness gaps, and recruit consented speakers for targeted improvement data.

Talk to our research team