Label: leaderboard-parity new models (cloning) — for merge into windows-cloning — ref chris_hemsworth_15s.wav
5 prompt(s) · one section per prompt · all models ranked by warm TTFA (fastest first) within each
Each prompt section shows every model's audio output, ordered by warm TTFA (fastest first). Click any audio player to hear that model's rendering.
Reference voice
Each model below was given this clip + transcript as the voice to imitate. Source: chris_hemsworth_15s.wav
Prompt 1
[en]"Open the browser and read my email."
Rank
Model
Device
TTFA warm
Audio
1
StyleTTS 2
cpu
129ms
2
StyleTTS 2
cuda
140ms
3
VibeVoice 7B
cuda
1.68s
4
Fish Speech 1.5
cuda
2.16s
5
Zonos v0.1
cuda
3.41s
6
Fish Speech 1.5
cpu
11.56s
7
Zonos v0.1
cpu
18.60s
Prompt 2
[en]"I'll start a new git branch, push the changes, and open a pull request when the tests pass."
Rank
Model
Device
TTFA warm
Audio
1
StyleTTS 2
cuda
173ms
2
StyleTTS 2
cpu
173ms
3
VibeVoice 7B
cuda
3.50s
4
Fish Speech 1.5
cuda
5.49s
5
Zonos v0.1
cuda
7.85s
6
Fish Speech 1.5
cpu
30.34s
7
Zonos v0.1
cpu
47.31s
Prompt 3
[en]"The Parakeet TDT zero point six billion parameter model achieves one point six nine percent word error rate on LibriSpeech test-clean, beating Whisper Large V3 at two point seven percent while running at over two thousand times realtime on a single GPU."
Rank
Model
Device
TTFA warm
Audio
1
StyleTTS 2
cuda
306ms
2
StyleTTS 2
cpu
340ms
3
VibeVoice 7B
cuda
9.65s
4
Fish Speech 1.5
cuda
20.28s
5
Zonos v0.1
cuda
23.91s
6
Fish Speech 1.5
cpu
107.15s
7
Zonos v0.1
cpu
157.87s
Prompt 4
[en]"Run pytest tests slash test underscore voice dot py with verbose flag and capture flag set to no."
Rank
Model
Device
TTFA warm
Audio
1
StyleTTS 2
cuda
228ms
2
StyleTTS 2
cpu
230ms
3
VibeVoice 7B
cuda
4.60s
4
Fish Speech 1.5
cuda
7.20s
5
Zonos v0.1
cuda
8.48s
6
Fish Speech 1.5
cpu
36.09s
7
Zonos v0.1
cpu
49.51s
Prompt 5
[fr]"Bonjour, je m'appelle Cicero et je vais vous aider avec votre code aujourd'hui."