Lifestyle

Hugging Face launches the Open TTS Leaderboard to compare open-source text-to-speech and voice cloning models with objective metrics

On September 30, 2026, Hugging Face announced the Open TTS Leaderboard. It ranks open-source text-to-speech models, which turn written text into spoken audio, by how accurately they speak, how fast they run and how closely a cloned voice matches the original. Here is how it works, what it cannot measure and why it matters to anyone who uses computer-generated voices.

About 6 min read

Hugging Face launches the Open TTS Leaderboard to compare open-source text-to-speech and voice cloning models with objective metrics
Image: Mokaair (Original editorial artwork)

What happened

On September 30, 2026, Hugging Face announced the Open TTS Leaderboard on its official blog. Text-to-speech (TTS) models are AI systems that read written text aloud in a synthetic voice. According to the company, the Hugging Face Hub hosted more than 8K TTS models as of that day, yet the ways of testing them remained fragmented and lacked standardization.

Hugging Face noted that existing arena-style leaderboards such as TTS Arena v2, Artificial Analysis and Voice Arena ask users to pick the better of two model outputs. Once enough votes have been collected, the arenas rank models with an Elo score, a rating calculated from those head-to-head choices. The company argues this approach cannot keep up with the pace of model releases. As an example, it said that as of September 30, 2026, only 16 of the 92 models on Artificial Analysis were open-weight, meaning their model files are publicly released so anyone can run them.

Hugging Face launches the Open TTS Leaderboard to compare open-source text-to-speech and voice cloning models with objective metrics
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

How the leaderboard scores models

  • Intelligibility: the generated speech is transcribed back into text by Qwen3 ASR, a speech recognition model, and the transcript is compared with the original text. The result is the word error rate (WER) or character error rate (CER), the share of words or characters that come out wrong.
  • Speed: RTFx, an inverse real-time factor that shows how much faster than real time a model produces audio, is measured for batched offline work on an H200 GPU, a high-end data-centre graphics processor. Time to first audio (TTFA) is measured on both an H200 GPU and a CPU.
  • Speaker similarity: for voice cloning, which copies a person's voice from a short reference recording, a score called SIM compares the generated speech with the reference clip. It uses WavLM speaker embeddings, numerical fingerprints of a voice, and measures how close they are with cosine similarity.

According to Hugging Face, the default view ranks models by their average WER across the English parts of two test sets, Seed TTS Eval and CV3 Eval (zero shot, meaning the models were not trained on these specific voices). In this view, hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead. For multilingual performance, the company singled out k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong performers. Chinese, Japanese and Korean are written in characters rather than space-separated words, so they are scored with CER.

The two ways of evaluating TTS models, as described by Hugging Face
AspectArena-style leaderboardsOpen TTS Leaderboard
Scoring methodUsers vote between two outputs; Elo scores are calculated from the votesObjective metrics such as WER/CER, RTFx, TTFA and SIM
Time needed to evaluate a modelA couple of weeks to collect votes (per Hugging Face)A couple of hours (per Hugging Face)
Open-source model coverage16 of 92 models on Artificial Analysis are open-weightCurrently focused on open-source models
Measures how natural a voice sounds?Yes, directly reflects human preferenceNo; Hugging Face says it does not replace human ratings

Listening, voice cloning and streaming speed

Turning on the "Voice cloning" option lets users compare only the models that support voice cloning and adds a SIM column. Hugging Face said models such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 showed improved average WER when given reference audio.

The "Listen" tab lets users hear and compare actual outputs from each model and submit feedback. Hugging Face asks voters to sign in with an HF account to help filter out spam and bots, and said voting data may be added to the leaderboard in the future.

The "Streaming" tab ranks models by TTFA, the wait between asking a model to speak and receiving the first playable audio. According to Hugging Face, each model is run on one request at a time (a batch size of 1) with the same 50 English CV3-Eval prompts and its default voice; the first 3 runs are discarded as warm-up and the median of the rest is reported. For non-streaming models, which cannot start playing until they finish, the time to generate the full utterance is measured. The company said kyutai/pocket-tts streams well on both GPU and CPU.

What it means for everyday users

For people who use voice assistants, audiobooks or read-aloud tools, leaderboards like this make it easier to compare open-source speech models on whether they pronounce text accurately, respond quickly and sound like the original voice. Still, Hugging Face itself stresses that these numbers do not show whether a voice sounds natural or pleasant, so how a voice actually sounds is best judged by listening.

Frequently asked questions

What is the Open TTS Leaderboard?

It is a leaderboard announced by Hugging Face on September 30, 2026. It ranks open-source multilingual text-to-speech and voice cloning models using objective metrics rather than listener votes.

Will it replace leaderboards based on human voting?

No. Hugging Face said it does not replace human preference rankings, because WER and speaker similarity cannot directly measure naturalness, expressiveness or listener preference. It can, however, help vote-based leaderboards choose which models to evaluate.

Does a model that does well in English necessarily do well in Chinese?

Not necessarily. Hugging Face said English performance does not necessarily carry over to other languages. The leaderboard therefore lets users switch rankings across multiple languages, with Chinese, Japanese and Korean scored by character error rate.

What is TTFA and why does it matter?

TTFA, or time to first audio, is the wait from sending a request until the first chunk of audio can be played. Hugging Face said this matters for interactive applications such as voice agents.

Will the evaluation code be released?

Hugging Face said it will open-source its evaluation scripts soon so the community can give feedback through GitHub Issues and pull requests. It has not announced a specific date.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle