Luke Oliff.

Eleven v4 Turbo Gets Voice Agents to Full Duplex

·Voice AI·5 min read·Luke Oliff
TL;DR

ElevenLabs released Eleven v4 and v4 Turbo on 28 September 2026. Turbo takes text while your LLM is still writing and sends audio back about 150 milliseconds later. In ElevenAgents a turn-taking model handles interruptions, so you get an agent you can talk over. The full duplex setup from my August post.

ElevenLabs released Eleven v4 and Eleven v4 Turbo on 28 September 2026. Turbo takes text while your LLM (the text model that writes the reply) is still writing it, and starts sending audio back about 150 milliseconds later. Inside ElevenAgents, a turn-taking model decides when the agent should stop and when it should keep going, so you get an agent you can talk over. ElevenLabs doesn’t call that full duplex. I do, because it’s the setup my August whitepaper asked for.

That post listed three problems with how voice agents are built. This one goes through them again with today’s release in hand, so it’s a follow on.

What ElevenLabs shipped

Two models came out together. Eleven v4 is for produced content, the audiobook and video work where quality matters more than speed. Eleven v4 Turbo is the fast version for agents. ElevenLabs says both are built on “an entirely new architecture” (a new way of building the model), and the product page describes something it calls context stitching, which keeps the same voice steady across a whole script.

The numbers from the product page:

  • More than 90 languages.
  • Instant Voice Clones from 10 seconds of audio.
  • Professional Voice Clones are back. They weren’t supported in v3.
  • One generation can be up to 10,000 characters.
  • Free plan: 10,000 credits a month, about 10 minutes of audio. Premium is $6 a month.

The blog post adds the test results. ElevenLabs says Artificial Analysis, an outside site that ranks AI models, puts v4 first for September 2026. It says about 75% of listeners picked v4 in blind tests, where the listener doesn’t know which model they’re hearing, and that v4 won 65% to 81% of head-to-head tests against Cartesia Sonic 3.6, Inworld TTS-2 and two of Google’s Gemini voices. Nobody outside ElevenLabs has published a measurement yet.

TechCrunch has the business side. Language support was 70 in v3. Revenue run rate, meaning this year’s pace rather than money already in, is past $600 million, up from about $330 million at the start of the year, and more than 55% of it comes from big companies.

What streaming both ways gives you

The old way to connect text-to-speech to an agent was to wait for the LLM to finish a sentence, send the sentence, then wait for the audio. Every sentence waited for the full trip out and back.

With v4 Turbo you push words as the LLM produces them and audio comes back before the sentence is finished. ElevenLabs’ words: “Bidirectional streaming, built for agent loops.” Latency, the wait before you hear anything, is about 150 milliseconds to the first sound. The model’s own processing time is about 100 milliseconds. Both are medians, the middle result with half faster and half slower. ElevenLabs’ comparison table puts Cartesia Sonic 3.6 at 262 milliseconds and OpenAI’s GPT-4o mini TTS at 814.

Is Eleven v4 Turbo full duplex? The three problems from August

The whitepaper described a voice agent as a chain of three parts, each waiting for the one before it to finish: speech-to-text, then the LLM, then text-to-speech. I called that chain the cascade. It broke the trouble with it into three problems.

The first was latency adding up at every step. Speech-to-text waits for enough audio, the LLM waits for the full transcript, and text-to-speech waits for the full text. v4 Turbo removes the last wait, and streaming both ways means it no longer waits for a full sentence either.

The second was that the model knows what it said but not what the user heard. The bigger the gap between what has been generated and what has been played, the more there is to sort out after an interruption. At 150 milliseconds to first speech, with audio going out as it’s made, that gap is now tiny.

Interruptions are the third problem, and the one I care about most. People make noises while they listen: “mm”, “yeah”, “uh-huh”, a laugh, a cough. On a call with a person those noises mean carry on, but a plain voice detector hears them as the user talking and stops the agent mid-sentence. ElevenAgents is built from four parts: a speech-to-text model, your LLM, a text-to-speech model (v4 Turbo now, according to the v4 blog post) and what the docs call “a proprietary turn-taking model that handles conversation timing”. Proprietary means their own, not public. A model whose only job is timing should be able to learn that a sound is not always a turn, and not always a change of subject. The docs don’t say what it’s trained on, and they list the interruption and turn-taking settings you can change.

What to do with it today

If you build voice agents, use Turbo for the speaking part, then make sure the rest of the chain, speech-to-text and the LLM, adds no more than another 150 milliseconds.

Then measure the three things the whitepaper asked for: the time from the user going quiet to the agent starting to speak, the time it takes to recover from an interruption, and how often the agent carries on talking after the user has started. If you’re on ElevenAgents, the turn-taking settings are where you change the last two.

One more thing from the agents docs: they still say the text-to-speech there covers 70 plus languages, while the v4 pages say 90 plus. Expect that page to catch up.