Open TTS Leaderboard: How Objective Metrics Are Finally Catching Up to Voice AI

Open TTS Leaderboard: How Objective Metrics Are Finally Catching Up to Voice AI

Text-to-speech has become a critical component of modern AI applications, from voice agents to accessibility tools. Yet despite the explosion of TTS models over the past year, we’ve been stuck with a broken evaluation framework. Arena-based leaderboards like Voice Arena and Artificial Analysis remain the gold standard for assessing quality, but they can’t scale fast enough to keep pace with releases. I’ve watched countless promising open-source models get overlooked simply because they lack the infrastructure burden of deployment that commercial APIs have. That’s starting to change.

The Open TTS Leaderboard represents a shift in how we think about model evaluation. Instead of relying purely on human preference votes (which take weeks to accumulate), the leaderboard uses objective metrics to evaluate models on complementary performance dimensions. The magic: evaluation drops from a couple weeks down to a couple hours. For a field moving this fast, that’s not incremental improvement, that’s transformative.

The Fragmentation Problem Nobody Wanted to Solve

Human preference remains the ultimate arbiter of quality. Metrics like MOS (Mean Opinion Score) and MUSHRA are still the academic standard. But here’s the uncomfortable truth: arena-based voting doesn’t scale, and it has structural biases baked in. Of 92 models on Artificial Analysis, only 16 are open-weights. The reason isn’t that open-source TTS is worse, it’s logistics. Adding an API model requires an API key. Hosting and serving an open model requires infrastructure investment and operational overhead that volunteer maintainers simply can’t afford.

There’s also the voter consistency problem. Even if you could scale arena voting, there’s no guarantee the same person evaluates models by the same criteria over time. As Heraclitus said, “a man cannot step into the same river twice.” Our preferences drift. Our understanding of what “better” means evolves. This inconsistency compounds when comparing models across weeks or months.

Objective metrics sidestep some of these challenges. The Open TTS Leaderboard ranks models using ASR-based Word Error Rate (WER) on datasets like Seed TTS Eval and CV3 Eval, combined with speaker similarity measurements to capture voice identity preservation. Neither directly measures naturalness or listener preference, but they provide reproducible, scalable proxies for quality dimensions that matter.

What This Means for Developers

I’m particularly interested in what this enables for practitioners. You can now compare TTS models across multiple dimensions in hours instead of weeks. The streaming comparison tab is especially valuable: it ranks models by TTFA (Time To First Audio), which measures latency from request to first audio chunk. For voice agents and interactive applications, this is often more important than overall quality. Seeing that kyutai/pocket-tts performs well on both GPU and CPU is exactly the kind of practical insight developers need when choosing a model for production.

The multilingual support also matters more than it initially appears. English performance doesn’t translate to other languages, and the leaderboard makes this explicit. You can toggle between languages and see which models genuinely handle multilingual inference well (spoiler: k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM’s offerings lead here). This prevents the common mistake of selecting a model based on English benchmarks only to discover it’s mediocre in your target language.

Voice cloning support introduces another layer. Some models like bosonai/higgs-tts improve substantially when given reference audio, which could significantly impact how you architect your application. The leaderboard makes these tradeoffs visible rather than requiring custom benchmarking.

The Listen Tab Changes Everything

What really stands out is the Listen tab. Most leaderboards are numbers in tables. This one lets you actually hear what these numbers represent. You can cherry-pick specific models or compare random selections across datasets. You can even vote on outputs, feeding community preference data back into the leaderboard itself. This is how you bridge the gap between objective metrics and human intuition.

This approach also acknowledges a critical limitation of purely metric-driven evaluation: they’re incomplete. WER doesn’t capture expressiveness. Speaker similarity doesn’t measure emotional nuance. The Listen tab isn’t a Band-Aid on this gap, it’s a deliberate design choice to keep humans in the loop while leveraging metrics for scale.

I think the most significant aspect is that the maintainers are explicitly inviting community feedback. They plan to open-source the evaluation scripts, similar to the Open ASR Leaderboard. This is how evaluation frameworks evolve. They become community-driven rather than beholden to whoever has the resources to operate them.

The Open TTS Leaderboard won’t replace arena-based voting for capturing true human preference. But it might finally democratize access to good evaluation data, making it possible for smaller teams and open-source maintainers to understand where their models stand without waiting weeks for voting results. What happens when evaluation infrastructure stops being a bottleneck for model development?

Read Next