5 prompt(s) · one section per prompt · all models ranked by warm TTFA (fastest first) within each
Each prompt section shows every model's audio output, ordered by warm TTFA (fastest first). Click any audio player to hear that model's rendering.
Prompt 1
[en]"Open the browser and read my email."
Rank
Model
Device
TTFA warm
Audio
1
F5-TTS v1
cuda
728ms
2
Dia 1.6B-0626
cuda
1.30s
3
IndexTTS-2
cpu
2.09s
4
MOSS-TTS-Nano
cuda
2.28s
5
IndexTTS-2
cuda
2.57s
6
MOSS-TTS-Nano
cpu
2.61s
7
MOSS-TTS v1.0
cuda
2.99s
8
F5-TTS v1
cpu
50.19s
Prompt 2
[en]"I'll start a new git branch, push the changes, and open a pull request when the tests pass."
Rank
Model
Device
TTFA warm
Audio
1
F5-TTS v1
cuda
914ms
2
MOSS-TTS-Nano
cuda
3.38s
3
MOSS-TTS v1.0
cuda
3.52s
4
MOSS-TTS-Nano
cpu
3.97s
5
IndexTTS-2
cpu
4.01s
6
IndexTTS-2
cuda
4.12s
7
Dia 1.6B-0626
cuda
8.34s
8
F5-TTS v1
cpu
64.77s
Prompt 3
[en]"The Parakeet TDT zero point six billion parameter model achieves one point six nine percent word error rate on LibriSpeech test-clean, beating Whisper Large V3 at two point seven percent while running at over two thousand times realtime on a single GPU."
Rank
Model
Device
TTFA warm
Audio
1
F5-TTS v1
cuda
1.71s
2
MOSS-TTS v1.0
cuda
8.96s
3
IndexTTS-2
cpu
9.05s
4
MOSS-TTS-Nano
cuda
9.17s
5
MOSS-TTS-Nano
cpu
10.40s
6
IndexTTS-2
cuda
10.75s
7
Dia 1.6B-0626
cuda
24.69s
8
F5-TTS v1
cpu
103.89s
Prompt 4
[en]"Run pytest tests slash test underscore voice dot py with verbose flag and capture flag set to no."
Rank
Model
Device
TTFA warm
Audio
1
F5-TTS v1
cuda
870ms
2
MOSS-TTS-Nano
cuda
3.69s
3
IndexTTS-2
cpu
4.32s
4
MOSS-TTS-Nano
cpu
4.58s
5
MOSS-TTS v1.0
cuda
4.69s
6
IndexTTS-2
cuda
5.19s
7
Dia 1.6B-0626
cuda
11.47s
8
F5-TTS v1
cpu
65.87s
Prompt 5
[fr]"Bonjour, je m'appelle Cicero et je vais vous aider avec votre code aujourd'hui."