Luke Oliff.

Open Weights Are Eating AI: Here Are the 7 Best Open TTS Models and the Gap Nobody Mentions

·Voice AI·6 min read·Luke Oliff

This is the biggest open-weight week AI has ever had: DeepSeek V4’s stable release lands July 24 and Moonshot’s Kimi K3 weights go free on July 27, days after K3 topped a major coding leaderboard against closed frontier models. In text, open weights have caught the leaders. So it’s a fair moment to ask what open weights buy you in speech, and the honest answer is: less than you’d hope. The best open-weight TTS model on the blind-vote Speech Arena sits around 1,118 Elo. The closed leaders sit at 1,236 and 1,234. That’s a gap of well over 100 points in a market where the top five closed models are separated by about 30.

Still, “behind the frontier” and “useless” are different things, and there are real reasons to run your own speech stack. Here are the seven open-weight TTS models worth knowing in July 2026, ranked by their blind-listener Elo, with the caveats attached.

Disclosure before the list: I work at Speechify on the SpeechifyAI API platform, which sells hosted TTS, so I’m structurally biased toward the “just use an API” conclusion. I’ve tried to let the numbers argue instead of me.

The 7 best open-weight TTS models by arena Elo

  • Step Audio EditX (StepFun), Elo 1,118. The best open-weight voice quality money doesn’t have to buy. Worth noting what that number means inside StepFun’s own catalogue: their hosted StepAudio 2.5 sits at 1,175 on the same board and lists at $85 per million characters. The open model gives up 57 Elo against its own commercial sibling and costs you only GPUs.

  • Fish Audio S2 Pro, Elo 1,110. Fish has quietly become the default recommendation in the self-hosting guides, and the arena score backs that up. Effective self-host cost works out around $5 per million characters once you price the GPU time, per the voice-agent cost surveys.

  • Voxtral TTS (Mistral), Elo 1,077. The most interesting engineering in the list. A 4B-parameter model that Mistral shipped in March 2026 with 70 to 90ms time-to-first-audio and voice cloning from 3 to 5 seconds of reference audio, across 9 languages. Mistral’s own blind evaluation put it ahead of ElevenLabs Flash v2.5 in 68.4% of cases, which is a vendor benchmark and should be read as one, but the arena Elo is independent and respectable.

  • Kokoro 82M v1.0, Elo 1,060. The one I keep recommending to hobbyists. At 82 million parameters it runs on hardware that embarrasses the rest of this list, and hosted versions cost $0.65 per million characters. I said in my leaderboard piece that cheap and good are different lists. Kokoro is the strongest argument that they at least rhyme.

  • Maya1, Elo 1,053. A newer entry that’s climbed steadily. Thin documentation trail so far; treat the score as the main signal.

  • Magpie-Multilingual, Elo 1,048. Does what the name says. If you need broad language coverage without an API bill, it’s the open-weight option built for that job, and it holds a respectable arena score while doing it.

  • Chatterbox, Elo 1,011. MIT-licensed, which matters: most of this list ships under custom or research licences that need legal review before commercial use. Its “beat ElevenLabs in a blind test” marketing claim predates its arena score settling 160+ points below ElevenLabs’ current models, so file that one under enthusiasm.

An honourable mention that isn’t arena-ranked yet: Qwen3-TTS, Apache-2.0, in 0.6B and 1.7B sizes, now the default TTS in Hugging Face’s speech-to-speech pipeline. Not to be confused with Qwen-Audio-3.0-TTS, the hosted commercial family whose Plus tier tops the whole leaderboard. I covered that naming mess [link: Qwen-Audio-3.0-TTS Flash coverage] this week.

Why is open-weight TTS so far behind open-weight text?

My theory, held loosely: speech quality at the top is now won with enormous amounts of curated, licensed, expressive audio data and RLHF-style listener feedback loops, and that’s exactly the input that’s hard to give away with a weights file. Kimi K3 could catch closed coding models because code and text are abundant and self-verifying. Natural prosody isn’t. The closed labs are also iterating monthly right now (the top of the arena changed hands twice in July alone), so the target moves faster than open projects release.

The gap is audible, too, which is the part the Elo numbers understate. A 30-point spread inside the closed top five is statistical noise between voices listeners can’t reliably rank. A 118-point spread is not. Listeners hear it.

When self-hosting TTS still wins

Four cases, honestly held. Privacy and compliance, where audio can’t leave your infrastructure and hosted-only models (including the current #1) are non-starters. Edge and offline deployment, where Kokoro’s 82M footprint does things no API can. Unmetered tinkering, because a weights file never sends you an invoice. And genuine cost wins at massive scale, though run the numbers before assuming this one: effective self-host costs land around $1 to $5 per million characters in GPU time before you pay an engineer to keep it up, and top-of-board hosted quality currently starts at $10 per million ($6 at volume) with someone else carrying the pager. The spreadsheet gap between “free” and the cheapest serious API has never been thinner, and that’s the closed market’s doing, not the open one’s.

FAQ

What is the best open-weight TTS model?

As of July 2026, Step Audio EditX by StepFun is the highest-ranked open-weight TTS model on the blind-vote Speech Arena, sitting at 1,118 Elo. Fish Audio S2 Pro is a close second at 1,110 Elo and is widely recommended for self-hosting.

How do open-weight TTS models compare to closed APIs?

Open-weight TTS models still trail the best closed APIs significantly. The top open model sits around 1,118 Elo, while the top closed models (like Qwen-Audio-3.0-TTS-Plus and Speechify’s Simba 3.2) sit around 1,236 and 1,234 Elo. This 100+ point gap represents a clearly audible difference in naturalness and prosody.

Is it cheaper to self-host an open-weight TTS model?

Yes, but the margin is shrinking. Self-hosting models like Fish Audio S2 Pro costs roughly $1 to $5 per million characters in GPU time (excluding engineering and infrastructure maintenance costs). Meanwhile, top-tier hosted APIs have dropped in price dramatically. For instance, Speechify’s Simba 3.2 delivers industry-leading quality starting at $10 per million characters ($6 at volume) with zero infrastructure overhead, making the build-vs-buy calculation heavily favor buying for production workloads.

If none of those four describe you, the boring conclusion stands: the arena’s price column already collapsed, and you can rent better-than-open quality for less than your GPU idle time costs. The interesting question is whether this week’s text-model shock eventually reaches speech. If a lab open-weights something within 30 Elo of the leaders, I’ll write the retraction happily. This week made that feel less hypothetical than it did in June.