Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text to Speech and Voice Cloning A leaderboard for open source and multilingual TTS Open source text to speech (TTS) d...
By AI Engineering Team
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
A leaderboard for open-source and multilingual TTS
Open-source text-to-speech (TTS) development has accelerated rapidly. As of September 30, 2026, more than 8,000 TTS models were available on the Hugging Face Hub. Evaluation has not advanced at the same pace, and remains fragmented across different datasets, metrics, and testing methods.
Human preference measures such as MOS and MUSHRA remain the standard for judging speech quality. Several arena-based leaderboards provide useful reference points:
These services present users with outputs from two TTS models and ask them to select the better result. After enough votes are collected, an Elo score is calculated, generally using the Bradley-Terry model.
Human preference remains the ultimate measure, but arena-based testing cannot easily keep pace with the number of new TTS releases. This may contribute to the limited representation of open-source models. On September 30, 2026, only 16 of the 92 models listed by Artificial Analysis were open-weight models, with a similar imbalance on Voice Arena.
There are practical reasons for this difference. Adding an API-based model generally requires an API key, while an open model must be hosted and served by the leaderboard operator. Commercial providers may also have greater incentives to seek placement than open-source developers. Arena testing also depends on voter consistency, which cannot be guaranteed over long periods because people may apply changing criteria when judging speech.
The Open TTS Leaderboard addresses scalability by using objective metrics across several performance dimensions:
- Intelligibility: Word error rate (WER) and character error rate (CER) are calculated between the prompt and the transcript of the generated audio. Transcription uses Qwen3 ASR, the highest-ranked open-source model on the Open ASR Leaderboard.
- Speed: Inverse real-time factor (RTFx) measures batched offline inference on an H200 GPU. Time-to-first-audio (TTFA) measures streaming latency with batch size 1 on an H200 GPU and CPU.
- Speaker similarity: Cosine similarity, or SIM, is calculated between WavLM speaker embeddings for the generated audio and the reference clip.
Using objective metrics reduces evaluation time from roughly a couple of weeks for vote collection to a couple of hours. The leaderboard does not replace human preference testing. ASR-based WER is a proxy for intelligibility, and speaker similarity estimates preservation of voice identity. Neither metric directly measures naturalness, expressiveness, or listener preference. The results can, however, help identify models for future voting-based evaluations.
Multilingual and voice-cloning evaluation
In the default leaderboard view, models are ranked by macro-average WER on the English splits of Seed TTS Eval, described in its paper, and CV3 Eval, evaluated in zero-shot mode and described in its paper.
hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro lead the English WER results when the two splits are averaged. Pareto plots show which models balance WER, batched inference speed measured by RTFx, and model size.
English results do not necessarily predict performance in other languages. The leaderboard allows users to select multiple languages. Seed TTS Eval provides audio for English and Chinese, while the remaining language results come from zero-shot CV3 Eval. Because Chinese, Japanese, and Korean are character-based languages, CER is reported for those languages. The cross-language “Average WER” is a macro-average across languages.
k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are among the stronger multilingual models in the evaluation.
Selecting Voice cloning limits the comparison to models that support voice cloning for the chosen languages. The table then includes a SIM column for speaker similarity, along with Pareto plots showing the tradeoffs between SIM, batched inference speed, and model size.
For some models, average WER improves when a reference recording is supplied for voice cloning. Examples include bosonai/higgs-tts-3-4b and openbmb/VoxCPM2.
Listening to and comparing TTS outputs
Metrics provide only part of the picture, so the leaderboard includes a Listen tab for direct comparison of generated audio.
Users can select:
- A language and dataset
- Whether to use voice cloning
- Specific models, or a random selection of models
The tab provides a way to examine outputs from multiple systems alongside the metric results. Users can also submit feedback on generated samples. As more community votes are collected, that information may be incorporated into the leaderboard.
Streaming performance
The Streaming tab compares how quickly models begin producing playable audio. Models are ranked by time-to-first-audio, or TTFA, which measures the delay between sending a request and receiving audio that can be played. This is particularly relevant to voice agents and other interactive applications.
For models with a streaming API, TTFA is the time until the first audio chunk arrives. For non-streaming models, it is the time required to generate the complete utterance, because playback cannot begin earlier.
Each model processes one audio request at a time, using batch size 1, the same 50 English prompts from CV3-Eval, the same hardware, and the model's default voice. The first three runs are discarded as warm-up measurements, and the reported TTFA is the median of the remaining runs.
The default results use an H200 GPU. CPU results are also available for a smaller, growing set of models. kyutai/pocket-tts performs well for streaming on both GPU and CPU.
Conclusion
The Open TTS Leaderboard is designed to provide faster, more consistent evaluation as new TTS models are released. Its initial focus is on:
- Open-source models, including systems that are often missing from arena-based evaluations.
- Multilingual performance, because English results are not a reliable proxy for other languages.
The evaluation scripts are planned for release as open source, following the approach used by the Open ASR Leaderboard repository.