Luke Oliff.

Qwen-Audio-3.0-TTS Plus Review: What Taking #1 Actually Bought Alibaba

·Voice AI·7 min read·Luke Oliff

Alibaba’s Qwen-Audio-3.0-TTS-Plus is the new #1 on the Artificial Analysis Speech Arena leaderboard, with an Elo of 1,236 (±17) from 1,305 blind listener votes. That puts it two points ahead of Simba 3.2 at 1,234 (±17), with confidence intervals that overlap almost entirely. It costs $27.6 per million characters and generates around 16 characters per second. So the new best TTS model in the world is in a statistical tie with the model below it, at 2.76x the price and roughly half the generation speed.

Disclosure before anything else: I work at Speechify, and Simba 3.2 is our model on the SpeechifyAI platform. The arena votes are blind and mine doesn’t count, but read this knowing where my salary comes from. I wrote up the full top ten [link: in last week’s leaderboard breakdown] and said I’d keep watching as new models dropped. Alibaba made that happen faster than I expected.

A horizontal bar chart showing the Speech Arena top 5 TTS models with their Elo scores and 95% confidence intervals. Qwen-Audio-3.0-TTS-Plus at 1236 and Simba 3.2 at 1234 have heavily overlapping confidence intervals, shown by a shaded region. The Speech Arena top five with 95% confidence intervals, July 19, 2026. The shaded band is where the two leaders’ intervals overlap: almost everywhere. The error bars are the review.

Is Qwen-Audio-3.0-TTS-Plus actually the best TTS model?

By the only fair public measure we have, yes, narrowly. The Speech Arena works on blind A/B votes: listeners hear two unlabeled samples from the same text and pick the more natural one, and those wins feed an Elo rating. No marketing budget can vote. Qwen-Audio-3.0-TTS-Plus climbed to the top within days of entering, past Gemini 3.1 Flash TTS (1,214) and Cartesia’s Sonic 3.5 (1,207), and Artificial Analysis credits it with noticeably natural, contextually appropriate intonation.

But a two-point lead inside a ±17 confidence interval is not a result, it’s a coin still spinning. With 1,305 arena appearances against Simba’s 1,275, both models are early in their vote accumulation and the gap between them is far smaller than the uncertainty around either number. The honest reading is that there are now two models sharing the top of the board, and the ranking between them could flip on any given week of votes.

And I’d go further than that. At this altitude on the board, I’m not sure “best” is a thing human ears can detect anymore. The top five sit within 30 Elo of each other with confidence intervals of ±13 to ±18, which means a listener in a blind test is often not hearing a better voice, they’re hearing a different voice, and voting for the timbre they happen to prefer. Quality at the top of this market has more or less peaked. The differences that remain are voice character, which is taste, and then the two things the leaderboard prints in the columns nobody screenshots: price and speed.

What Alibaba has genuinely done is put a Chinese lab at the top of a leaderboard that US companies had to themselves a month ago. That’s the headline, and it deserves to be.

What does it cost, and is that justified?

$27.6 per million characters puts Qwen-Audio-3.0-TTS-Plus in an odd middle band. It’s nearly three times Simba 3.2 ($10, and $6 on the Scale tier), yet a bargain next to Sonic 3.5 ($49), StepAudio 2.5 ($85), or MiniMax and ElevenLabs at $100. For quality that the arena says is indistinguishable from the $10 model, you’re paying $17.6 extra per million characters for… the tiebreaker vote, I suppose.

Speed is where the case gets harder to make. The model generates about 16 characters per second. Spoken English runs around 12 to 14 characters per second, so generation barely outpaces playback, and that’s before network overhead. Simba 3.2 generates at 30.2 chars/sec and Sonic 3.5 at 120. For batch narration jobs that’s an inconvenience; for real-time agent and streaming workloads it’s disqualifying, because your buffer never gets ahead of the listener. If Alibaba ships a faster variant (they usually do, the Qwen team iterates relentlessly), this criticism expires. Today it stands.

Where the price probably is justified: if your product serves Chinese-language users. The Qwen audio stack’s dialect and Mandarin coverage is deeper than anything the US providers offer, and the same week this model took #1, the Qwen team shipped a real-time model upgrade covering 16 dialects and 30 languages. If that’s your market, nothing else on the leaderboard is really a substitute, and $27.6 is fine.

Who should switch to it?

Nobody should switch on the Elo alone, and I’d say that even if the model it displaced wasn’t ours. A two-point gap inside overlapping confidence intervals will not be audible in your product. Once you accept that the top of the board is a quality plateau, the selection criteria flip: shortlist by price and latency first, then filter the arena by your category (the board splits out Assistants, Entertainment, Customer Service, and Knowledge Sharing) and let your own users pick between the two or three voices that survive the budget. Working top-down from the Elo column is how you end up paying $100 per million characters for a voice your listeners can’t distinguish from a $10 one.

What this release does change is the shape of the market. A month ago the top of the board was American, and now the #1 quality signal belongs to Alibaba at a mid-tier price. Enterprise buyers who can’t use a China-hosted API for compliance reasons will carry on as before. Everyone else just got another serious option, and the vendors charging $85 to $100 per million characters for quality below both leaders got another very bad week.

A scatter plot showing TTS model quality (Elo score) on the Y-axis versus price (USD per million characters) on the X-axis. Simba 3.2 is prominently highlighted in the top-left sweet spot quadrant. The lonely dots at $100 (ElevenLabs and MiniMax) are isolated on the right. Quality vs price across the priced top ten, July 19, 2026. Simba had the top-left quadrant to itself last week and that hasn’t changed; Qwen took the #1 rank without touching the price frontier. The lonely dot at $100 is doing a lot of work for MiniMax’s pricing team.

FAQ

What is Qwen-Audio-3.0-TTS-Plus?

It’s Alibaba’s latest text-to-speech model, released in July 2026 by the Qwen team. It currently ranks #1 on the Artificial Analysis Speech Arena leaderboard with an Elo of 1,236 from blind listener votes, priced at $27.6 per million characters via Alibaba’s API.

How does it compare to Simba 3.2?

They’re statistically tied on quality: 1,236 vs 1,234 Elo with ±17 confidence intervals that overlap almost completely, so blind listeners are effectively choosing between two different voices rather than a better one. The measurable differences are price (Simba is $10 per million characters against $27.6) and generation speed (Simba produces about 30.2 characters per second against roughly 16). For production workloads, Simba 3.2 is the clear choice given its massive advantages in speed and cost. Disclosure: Simba is built by Speechify, where I work.

Is 16 characters per second fast enough for real-time TTS?

Barely, and only in ideal conditions. English speech plays back at roughly 12 to 14 characters per second, so a 16 chars/sec model leaves almost no buffer headroom once you add network latency. Fine for pre-generated audio, risky for live conversational agents where the model must stay ahead of playback.

How often does the Speech Arena leaderboard change?

Constantly. Votes stream in around the clock and new models are added as they launch. Qwen-Audio-3.0-TTS-Plus reached #1 within days of being listed, and most of the current top ten shipped within the last six months. Check the live board rather than trusting any article’s snapshot, including this one.