Luke Oliff.

Google just cut the price of decent speech by 70%, quietly

·AI·5 min read·Luke Oliff

Google just cut the price of decent speech by 70%, quietly

By Luke Oliff | Jul 2026 | 5 min read


Gemini 3.5 Flash TTS is live at $6 per million output tokens, down from the $20 that Gemini 3.1 Flash TTS charged when it shipped in April 2026. That works out to roughly $0.54 for an hour of generated audio, or about $0.09 for a ten-minute narration. Google didn’t hold a launch event for this. There was no keynote, no benchmark montage. They just moved the number, and the number is the story.

I’ve been watching Google’s TTS pricing do this dance for a while, and the pattern is worth laying out properly because it tells you what Google thinks its position in this market actually is.

What changed between Gemini 3.1 and 3.5 Flash TTS?

The 3.1 release in April was the feature release. It brought 70+ languages, 200+ inline audio tags like [whispers] and [laughs], two-speaker support, and SynthID watermarking on everything it outputs, and it earned Google its highest leaderboard placement ever, currently sitting at 1,214 Elo on the Artificial Analysis Speech Arena. For about ten weeks it was arguably the best value at the top of the board.

The 3.5 release is the price release. Token output pricing dropped from $20 to $6 per million. Audio is metered at 25 tokens per second of generated speech: an hour of audio is 90,000 tokens, which cost $1.80 on 3.1 and costs $0.54 now.

For context, Gemini 2.5 Flash TTS charged $10 per million output tokens a year ago. So the sequence runs $10, then $20 for the big feature generation, then $6. Google raised the price when it had something new to sell and cut it the moment the market caught up. That’s not sloppy pricing, that’s a company telling you it intends to compete on cost.

What does $6 per million tokens mean against character pricing?

An hour of spoken English at a typical 150 words per minute is roughly 54,000 characters (six characters per word including spaces).

At that rate, per hour of audio:

  • Gemini 3.5 Flash TTS: ~$0.54
  • Simba 3.2 at $10 list: ~$0.54; at $6 Scale tier: ~$0.32
  • Qwen-Audio-3.0-TTS-Plus: ~$1.49
  • Cartesia’s Sonic 3.5 at $49/M chars: ~$2.65
  • ElevenLabs v3 at $100/M chars: ~$5.40

Google’s new price lands almost exactly on top of the cheapest high-quality character-priced models, to the cent in one case. Disclosure: I work at Speechify, on the SpeechifyAI platform side, and Simba is our model. When the biggest infrastructure company on earth prices its speech output to match you, that’s flattering and also mildly terrifying.

The remaining difference is the input charge. Gemini bills text input at $1 per million tokens on top of output, which adds about $0.014 to that hour. Rounding error for narration, but it compounds oddly for agent workloads where you resend context.

Is Gemini 3.5 Flash TTS worth switching to?

If you’re already on 3.1 Flash TTS, almost certainly, assuming quality holds. The feature set carried over and you’re being offered the same output for 30% of the price. Run your regression suite on pronunciation edge cases first (Google’s TTS models have historically wobbled on domain jargon when versions change), but this is the easy call.

If you’re on a character-priced provider, slow down. Token pricing has a budgeting problem: your finance team can count the characters in a script before you spend anything, but tokens are only knowable after generation. 25 tokens per second is predictable enough once audio length is fixed, yet audio length is exactly the thing you don’t control when a model decides to pause, or breathe, or take the [dramatic] tag seriously.

And the leaderboard question is genuinely open. 3.1 Flash TTS earned its 1,214 Elo over months of blind votes. 3.5 hasn’t been through that grinder yet. I lean confidence, given the 3.1 pedigree, but I’d want two weeks of arena data before moving a production narration pipeline.

The bigger picture is that top-tier speech pricing has fallen 70% inside a quarter at Google alone, and every provider still listing at 2024 rates is now visibly exposed. If your TTS contract predates this cut, requote it. This week, not next quarter.

FAQ

How much does Gemini 3.5 Flash TTS cost? $6 per million output tokens, plus $1 per million input tokens. Audio is metered at 25 tokens per second, so an hour of audio is ~$0.54. A ten-minute narration runs ~$0.09.

Is Gemini 3.5 Flash TTS better than Gemini 3.1 Flash TTS? On price, clearly: 70% cheaper per token. On quality, the evidence isn’t in yet. 3.1 holds 1,214 Elo from months of blind listener votes; 3.5 hasn’t accumulated enough arena appearances to compare.

How do token-priced and character-priced TTS APIs compare? Convert both to cost per hour of audio. At July 2026 rates, an hour of speech runs $0.32 to $0.54 at the cheap end of both schemes and over $5 at the expensive end.

What is the cheapest good TTS API in July 2026? Among models with top-ten leaderboard quality: Simba 3.2’s $6/M chars Scale tier ($0.32/hour) and Gemini 3.5 Flash TTS at $6/M output tokens ($0.54/hour).


Tags: AI, Text To Speech, Google, Deepmind, API