Lifestyle
Hugging Face launches the Open TTS Leaderboard to compare open-source text-to-speech and voice cloning models with objective metrics
On September 30, 2026, Hugging Face announced the Open TTS Leaderboard. It ranks open-source text-to-speech models, which turn written text into spoken audio, by how accurately they speak, how fast they run and how closely a cloned voice matches the original. Here is how it works, what it cannot measure and why it matters to anyone who uses computer-generated voices.
About 6 min read

What happened
On September 30, 2026, Hugging Face announced the Open TTS Leaderboard on its official blog. Text-to-speech (TTS) models are AI systems that read written text aloud in a synthetic voice. According to the company, the Hugging Face Hub hosted more than 8K TTS models as of that day, yet the ways of testing them remained fragmented and lacked standardization.
Hugging Face noted that existing arena-style leaderboards such as TTS Arena v2, Artificial Analysis and Voice Arena ask users to pick the better of two model outputs. Once enough votes have been collected, the arenas rank models with an Elo score, a rating calculated from those head-to-head choices. The company argues this approach cannot keep up with the pace of model releases. As an example, it said that as of September 30, 2026, only 16 of the 92 models on Artificial Analysis were open-weight, meaning their model files are publicly released so anyone can run them.
Read the full description
Sources are collected, independently checked, then reviewed by Jev.
How the leaderboard scores models
- Intelligibility: the generated speech is transcribed back into text by Qwen3 ASR, a speech recognition model, and the transcript is compared with the original text. The result is the word error rate (WER) or character error rate (CER), the share of words or characters that come out wrong.
- Speed: RTFx, an inverse real-time factor that shows how much faster than real time a model produces audio, is measured for batched offline work on an H200 GPU, a high-end data-centre graphics processor. Time to first audio (TTFA) is measured on both an H200 GPU and a CPU.
- Speaker similarity: for voice cloning, which copies a person's voice from a short reference recording, a score called SIM compares the generated speech with the reference clip. It uses WavLM speaker embeddings, numerical fingerprints of a voice, and measures how close they are with cosine similarity.
According to Hugging Face, the default view ranks models by their average WER across the English parts of two test sets, Seed TTS Eval and CV3 Eval (zero shot, meaning the models were not trained on these specific voices). In this view, hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead. For multilingual performance, the company singled out k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong performers. Chinese, Japanese and Korean are written in characters rather than space-separated words, so they are scored with CER.
| Aspect | Arena-style leaderboards | Open TTS Leaderboard |
|---|---|---|
| Scoring method | Users vote between two outputs; Elo scores are calculated from the votes | Objective metrics such as WER/CER, RTFx, TTFA and SIM |
| Time needed to evaluate a model | A couple of weeks to collect votes (per Hugging Face) | A couple of hours (per Hugging Face) |
| Open-source model coverage | 16 of 92 models on Artificial Analysis are open-weight | Currently focused on open-source models |
| Measures how natural a voice sounds? | Yes, directly reflects human preference | No; Hugging Face says it does not replace human ratings |
Listening, voice cloning and streaming speed
Turning on the "Voice cloning" option lets users compare only the models that support voice cloning and adds a SIM column. Hugging Face said models such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 showed improved average WER when given reference audio.
The "Listen" tab lets users hear and compare actual outputs from each model and submit feedback. Hugging Face asks voters to sign in with an HF account to help filter out spam and bots, and said voting data may be added to the leaderboard in the future.
The "Streaming" tab ranks models by TTFA, the wait between asking a model to speak and receiving the first playable audio. According to Hugging Face, each model is run on one request at a time (a batch size of 1) with the same 50 English CV3-Eval prompts and its default voice; the first 3 runs are discarded as warm-up and the median of the rest is reported. For non-streaming models, which cannot start playing until they finish, the time to generate the full utterance is measured. The company said kyutai/pocket-tts streams well on both GPU and CPU.
What it means for everyday users
For people who use voice assistants, audiobooks or read-aloud tools, leaderboards like this make it easier to compare open-source speech models on whether they pronounce text accurately, respond quickly and sound like the original voice. Still, Hugging Face itself stresses that these numbers do not show whether a voice sounds natural or pleasant, so how a voice actually sounds is best judged by listening.
Frequently asked questions
What is the Open TTS Leaderboard?
It is a leaderboard announced by Hugging Face on September 30, 2026. It ranks open-source multilingual text-to-speech and voice cloning models using objective metrics rather than listener votes.
Will it replace leaderboards based on human voting?
No. Hugging Face said it does not replace human preference rankings, because WER and speaker similarity cannot directly measure naturalness, expressiveness or listener preference. It can, however, help vote-based leaderboards choose which models to evaluate.
Does a model that does well in English necessarily do well in Chinese?
Not necessarily. Hugging Face said English performance does not necessarily carry over to other languages. The leaderboard therefore lets users switch rankings across multiple languages, with Chinese, Japanese and Korean scored by character error rate.
What is TTFA and why does it matter?
TTFA, or time to first audio, is the wait from sending a request until the first chunk of audio can be played. Hugging Face said this matters for interactive applications such as voice agents.
Will the evaluation code be released?
Hugging Face said it will open-source its evaluation scripts soon so the community can give feedback through GitHub Issues and pull requests. It has not announced a specific date.
Browse the latest news in this topic
Lifestyle
Google Cloud Launches Spanner Queues: Putting Message Queues Inside Database Transactions to Make AI Agents More Reliable
Google Cloud has announced the general availability of Spanner queues, which make message creation part of a database transaction. The aim is to stop AI agents' "state" and "actions" from falling out of sync. This article covers Google Cloud's claims, the main features, and what it means for general readers.
Lifestyle
GPT-6.1 Sol Launches: New Sol Version in the API, Codex and ChatGPT Work, Not in Chat
OpenAI launched GPT-6.1 Sol on September 29, 2026, with the API name gpt-6.1-sol. The launch rollout covers Codex and ChatGPT Work on Plus, Pro, Business, Enterprise and Edu (Enterprise and Edu need an administrator to enable it); Free and Go are not included at launch, and it is not in Chat (checked September 2026).
Lifestyle
Claude Sonnet 5.5 Launches: Same List Price as Sonnet 5, Available in the API, on Cloud Platforms and in Claude.ai
Anthropic launched Claude Sonnet 5.5 on September 28, 2026. API list prices are the same as Sonnet 5 ($2 per million input tokens, $10 per million output tokens). It is available in Claude.ai, the API and several cloud platforms, and higher-risk cybersecurity requests fall back to Sonnet 5 (checked September 2026).
Lifestyle
Claude Code Adds mods: TypeScript Functions in Plugins Change Its Behavior and Interface
On October 1, 2026, Anthropic introduced Claude Code mods: TypeScript or JavaScript functions that ship inside a plugin and run in the Claude Code process, and can rewrite prompts, manage tool calls and draw new interface. Anthropic says they work in the CLI and the desktop app and are not sandboxed; the documentation says v2.1.287 or later is required (checked October 2026).
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family