ListenSpeedScoresCapabilitiesπŸ—³ Vote β†—

Scores

Objective scores over the same 5 prompts. UTMOS = predicted naturalness (higher better); WER = ASR word-error rate vs the intended text β€” a failure-detector, not a fine ranking (lower better); SIM = speaker similarity to the cloned reference (chris_hemsworth_15s, higher better); Health = deterministic defect triage of the published clip (⚠ flags long internal silence / clipping / dead audio β€” a "go listen" cue, not a score). Switch Default / Cloning below; click any header to re-sort. Each score is the mean over the exact clips shown on Listen. Human votes are the preference ground truth; these objective metrics are backstops.
Voice:
Default voiceCloning

Default voice Β· naturalness + intelligibility

ModelSizeReleasedUTMOS ↑WER ↓Health
Chatterbox1.2BApr 20254.4230.061βœ“
Chatterbox Turbo744MDec 20254.2740.074βœ“
Coqui XTTS-v2750MOct 20234.0560.065βœ“
Dia 1.6B-06261.6BJun 20252.3870.317βœ“
dots.tts (soar)2BJun 20263.2270.059βœ“
DramaBox3.3BApr 20264.3040.093βœ“
F5-TTS v1330MOct 20244.0810.107βœ“
Higgs Audio v3 TTS4BJun 20264.3720.065βœ“
IndexTTS-21.5BJun 20254.2570.086βœ“
KittenTTS Nano 0.1<100MAug 20253.6650.093βœ“
Kokoro82MDec 20244.3020.065βœ“
LFM2.5-Audio 1.5B1.5BDec 20254.3830.066βœ“
LongCat-AudioDiT 1B1.42BMar 20264.0530.093βœ“
LongCat-AudioDiT 3.5B3.83BMar 20264.3410.080βœ“
LuxTTS123MJan 20263.5400.092βœ“
Magpie-TTS357MDec 20254.1990.087βœ“
Mars5-TTS1.2BJun 20243.5400.361βœ“
Maya13BOct 20254.4870.066βœ“
MeloTTS~52MFeb 20243.4980.058βœ“
MiraTTS0.5BDec 20253.8030.100βœ“
Miso TTS 8B8.2BMay 20264.2320.586βœ“
NeuTTS Air748MSep 20254.0030.117βœ“
NeuTTS Nano229MDec 20253.5720.066βœ“
OmniVoice~1BMar 20264.1040.040βœ“
Orpheus-TTS 3B3.3BMar 20254.0050.102βœ“
OuteTTS 1.0 1B1BApr 20254.3860.070βœ“
Parler-TTS Mini v1878MJun 20243.7570.149⚠ gap (3)
Piper~25MBJan 20234.0770.066βœ“
Pocket-TTS100MJan 20264.0970.054βœ“
Qwen3-TTS 1.7B (CUDA-graph)1.7BJan 20264.3230.065βœ“
Qwen3-TTS 1.7B Base1.7BJan 20264.2760.096βœ“
Scylla's Band~103MJul 20264.4440.081βœ“
Sesame CSM-1B1BMar 20254.1520.114βœ“
Soprano 1.1 80M80MJan 20264.1160.059βœ“
Step-Audio-EditX3BOct 20254.3990.044βœ“
StyleTTS 2~148MJun 20234.2590.158βœ“
Supertonic 399MMay 20264.1950.065βœ“
VibeVoice Realtime 0.5B0.5BDec 20254.0430.148βœ“
VoxCPM2 2B2BApr 20263.4800.023βœ“
Voxtral 4B TTS4BNov 20253.6920.081βœ“
Scored over the 5 bench prompts (thin β€” WER is a failure-detector, not a fine ranking). Checkpoints: UTMOS utmos22_strong (SpeechMOS), SIM canonical UniSpeech-SAT wavlm_large_finetune, WER Whisper-large-v3. Method follows seed-tts-eval. Human votes are the preference ground truth; these are objective backstops.