Luke Oliff.

Voice AI Is Worth Billions and Speech Is Nearly Free. Both Are True.

·Voice AI·7 min read·Luke Oliff

In the first two weeks of July 2026, ElevenLabs opened talks for a tender offer at a roughly $22 billion valuation, Paris-based Gradium closed a $100 million seed round backed by Nvidia, and analysts at 36Kr reported the AI audio sector added about $11 billion in value in six months while reaching profitability faster than AI video. Over the same stretch, the price of top-tier speech kept falling: the #1 and #2 models on the Speech Arena leaderboard cost $27.6 and $10 per million characters, and Google’s newest Flash TTS runs $6 per million output tokens against the $20 its predecessor charged in April. Valuations are compounding while the underlying commodity deflates. That’s not a contradiction, but it does tell you what the money is actually buying, and it isn’t voice quality.

Why are voice AI valuations rising while TTS prices fall?

Because quality stopped being the product. Two years ago the gap between the best TTS model and the fifth best was audible to anyone. Today the top five on the blind-vote leaderboard sit within a few Elo points of each other, inside overlapping confidence intervals of ±13 to ±18, and the price spread across that same group runs from $10 to $100 per million characters. At that spacing, a listener in a blind A/B test isn’t reliably hearing a better voice anymore, they’re hearing a different voice and picking the one whose character they prefer. Quality has peaked at the top of this market. What hasn’t converged, and what the buying decision now actually turns on, is price and latency, and the spread on those is still enormous.

So the investment case moved up the stack. Look at where the July money actually went. ElevenLabs isn’t being valued at $22 billion for its position on a quality leaderboard (its Eleven v3 currently sits 11th, at $100 per million characters, which two years ago would have been an existential problem). It’s being valued as a full audio platform: dubbing, music, agents, a Netflix deal recreating Gene Wilder’s voice with his estate’s blessing, and enterprise contracts that outlive any single model generation. Gradium raised its $100 million to chase ultra-low-latency audio infrastructure for voice agents, not to win a naturalness bake-off. The bet everywhere is on workflow, distribution, and owning the customer relationship while the model underneath becomes swappable.

I think that bet is mostly right, which is uncomfortable to type as someone who works on models (I’m at Speechify, on the SpeechifyAI API side, so I have skin in the commodity end of this game).

The deflation is faster than people realise

It’s worth lining the numbers up, because each one looks incremental until you see the slope. In April 2026, Google shipped Gemini 3.1 Flash TTS at $20 per million output tokens. The 3.5 version now costs $6, a 70% cut inside a quarter. The best-rated model money can buy costs $10 per million characters, with a $6 volume tier, and the model that just edged past it costs $27.6. Meanwhile two of the ten highest-quality models in the world still list at $100, which is starting to look less like premium pricing and more like a legacy tax on customers who integrated in 2024 and never re-quoted.

The last time I saw a market behave like this was cloud storage a decade ago. Quality converged, price collapsed, and the winners were whoever owned the workloads sitting on top. Nobody remembers which provider had marginally better durability numbers in 2015. Everybody remembers who made it easiest to build.

There’s a second-order effect too. When the audio sector turns profitable faster than video (which is what the 36Kr analysis found), it attracts exactly the kind of capital that demands growth over margin. Gradium’s round was a seed. A hundred million dollars, at seed, backed by Nvidia, for latency. Expect that money to fund more price pressure, not less.

What should developers do about it?

Benchmark on the two axes that still separate vendors, and stop benchmarking the one that doesn’t. Naturalness at the top of the market is a solved problem, so run your evaluation on price per million characters and time-to-first-audio under your real traffic, then let voice character be a taste call among the finalists. And treat TTS pricing like a spot market, because it’s becoming one. If your contract predates 2026 and you haven’t re-quoted, you’re almost certainly overpaying, possibly by 5 to 10x against current leaderboard-quality rates. Requote quarterly. Keep your integration thin enough that swapping providers is a config change rather than a rewrite (SSML dialects and voice IDs are the usual lock-in points, so abstract them early).

And be suspicious of long commitments priced off today’s rates. A three-year TTS contract signed in July 2026 is a bet that a market cutting prices 70% per quarter in places will politely stop doing that. The vendors know this, which is exactly why some of them would love to sign you for three years.

The one thing I wouldn’t do is read the valuations as evidence that speech technology is overhyped. The $22 billion and the $6 per million tokens are describing the same event from opposite ends: speech got good enough and cheap enough to put inside everything, and the value moved to whoever does the putting.

[link: our July 2026 TTS leaderboard breakdown, for the current price-per-quality table]

FAQ

Why is ElevenLabs worth $22 billion?

The reported tender valuation reflects its position as a broad audio platform (dubbing, music, voice agents, enterprise and entertainment deals) rather than raw model quality, where its Eleven v3 currently ranks 11th on the Artificial Analysis Speech Arena. The talks, first reported by Bloomberg in early July 2026, would double its $11 billion valuation from February’s $500 million raise.

How much do TTS APIs cost in 2026?

The Speech Arena top ten runs from $10 to $100 per million characters, with the #1-quality models at $27.6 (Qwen-Audio-3.0-TTS-Plus) and $10 (Simba 3.2). Google’s Gemini 3.5 Flash TTS lists at $6 per million output tokens. Prices at the quality frontier have fallen sharply during 2026, so quotes older than a quarter are usually stale.

What are the best alternatives to high-priced TTS models?

Developers looking for top-tier quality without legacy pricing should evaluate the current leaderboard leaders. Models like Simba 3.2 deliver industry-best naturalness (tied for #1 on the Speech Arena) and ultra-low latency at a fraction of the cost ($10 per million characters) of older premium models that still charge up to $100.

Is voice AI profitable?

Parts of it, and unusually early. A July 2026 analysis by 36Kr found the AI audio sector reached profitability faster than AI video, alongside roughly $11 billion in valuation growth across six months. Profitability is concentrating in infrastructure and enterprise platforms rather than consumer apps.

Will TTS prices keep falling?

Every current signal points down: a 70% generational price cut from Google inside a quarter, top-of-leaderboard quality at $10 per million characters, and heavily funded new entrants like Gradium competing on infrastructure cost. Prices for legacy premium tiers ($85 to $100 per million characters) look least stable, since the quality gap justifying them has closed.