Luke Oliff.

Voice AI.

72 posts

Full-Duplex Voice AI Needs a New ArchitectureOriginalThe PACE paper exposes a fundamental flaw in LLM voice dialogue: the model sees context the user never heard. Full-duplex voice needs a new architecture, not just faster TTS.Voice AI
Sunday roundup: six posts from a week in voice AIOriginalSix posts this week. EU watermarking, emotion prompts, open source voice agents, AI coding one year on, a raccoon heist game, and the AI industry roundup.Voice AI
Open Source Voice Agents Get Real-Time Speech in Hermes v0.20.0OriginalHermes Agent v0.20.0 brings streaming TTS, barge-in, on-device wake words, and pluggable STT/TTS to open source voice agents. Released August 3, 2026.Voice AI
Voice Emotion Control Moves From SSML to PromptsOriginalVoice emotion control is moving to natural language prompts. Kakao's Kanana-o scores 94.50 on the Korean InstructTTSEval benchmark.Voice AI
EU AI Act Voice Watermarking: What TTS Builders Must KnowOriginalEU AI Act voice watermarking rules took effect August 2, 2026. What TTS providers and voice developers must do to stay compliant.Voice AI
Smallest.ai Raises $13M to Split Voice Agents in TwoOriginalSmallest.ai closed a $13M Series A for Voice 4.0 and Hydra. Their bet is that fast real-time agents need small speech models, not big LLMs.Voice AI
GPT-Transcribe Makes Context the New ASR FeatureOriginalOpenAI's GPT-Transcribe launched July 29, 2026 with prompt, keyword, and language hints. Free-form context lifted accuracy from 38.5% to 44.6%.Speech-to-Text
Grok Voice 2.0 Ships With a Quiet 60% Price RiseOriginalGrok Voice Think Fast 2.0 launched July 29, 2026 at $0.08/min, up from $0.05. On August 5 the grok-voice-latest alias migrates you automatically.Voice AI
Fish Audio's $52M Seed: Open Weights Got Them HereOriginalFish Audio raised a $52M seed on July 28, 2026 with $21M ARR and 8M users. Its new S2.1 Pro model is closed, API-only. What that shift means for TTS.TTS
OpenAI Presence: Voice Agents You Can't Self-ServeOriginalOpenAI Presence deploys enterprise voice and chat agents, but only via OpenAI's own engineers. No API, no pricing, no self-serve. What that means.Voice AI
What a Week at SpeechifyOriginalSimba went multilingual, streaming got timestamps, voice agents learned to switch languages mid-call. Qwen took the crown. MCP broke everything.Voice AI
MAI-Voice-Flash, Opus 5, and a Red LineOriginalMicrosoft MAI-Voice-2-Flash enters TTS at $15/M. Anthropic ships Opus 5 at half Fable's cost. OpenAI faces red-line questions after Hugging Face.Voice AI
The 7 Best Open TTS Models as Open Weights Eat AIOriginalDeepSeek V4 and Kimi K3 made this the biggest open-weight week AI has had. Here are the 7 best open TTS models, and the gap nobody mentions.Voice AI
Fortnite Is About to Be Gaming's Biggest TTS DeploymentOriginalEpic Games gives 36 Fortnite characters AI voices on July 30, 2026, making it the biggest real-time TTS deployment gaming has seen.Voice AI
Five Voice AI Stories from July 2026 That MatterOriginalMid-July 2026 gave voice AI a new leaderboard king, a 70% Google price cut, a $22 billion valuation rumour, and an AI Gene Wilder.Voice AI
Qwen-Audio-3.0-TTS Plus Review: What #1 Bought AlibabaOriginalAlibaba's Qwen-Audio-3.0-TTS-Plus tops the Artificial Analysis Speech Arena at 1,236 Elo from 1,305 blind votes. A review of what that buys.Voice AI
Synthetic Voice Rules Arrived from Three DirectionsOriginalIn one week, synthetic voices got rules at three layers of the stack: a TikTok Shop ban, platform policy, and government regulation.Voice AI
Brussels Just Gave Voice Assistants the Keys to AndroidOriginalThe EU's Digital Markets Act decisions of July 16, 2026 order Google to open Android to rival AI assistants, with certified voice access.Voice AI
AMD Just Put Text-to-Speech in the Local AI Stack by DefaultOriginalAMD's Lemonade 11.0 puts text-to-speech in the local AI stack by default: an OpenMOSS backend, voice cloning, and a dedicated TTS panel.Voice AI
Listeners Voted and the AI Voices WonOriginalTwo mid-July 2026 studies found ordinary listeners can't reliably tell AI voices from human ones, and sometimes prefer the synthetic option.Voice AI
TTS Quality Has No Single Number YetOriginalTTS quality has no single metric. MOS scores cluster, Elo only tests short English clips, and TTS WER measures a different thing than STT. Here is how to actually evaluate.TTS
Voice AI Is Worth Billions While Speech Is Nearly FreeOriginalIn two weeks of July 2026, ElevenLabs talked a $22 billion tender offer and Gradium raised $100 million while speech itself got nearly free.Voice AI
Gene Wilder's AI Voice: The Paperwork Is the StoryOriginalNetflix is using an AI recreation of Gene Wilder’s voice in “Wonka’s The Golden Ticket,” a nine-episode reality competition [premiering September 23,…Voice AI
GPT-Live Failed Its First Viral Test in 12 SecondsOriginalOpenAI's GPT-Live launched July 8 and the internet found its weak spot in hours. TikToker Husk broke it with a spelling test, and the full-duplex interruptions are already a meme.Voice AI
OpenAI's Best Voice Model Is Locked Away from DevelopersOriginalOpenAI's GPT-Live-1 full-duplex voice models replaced Advanced Voice Mode in ChatGPT on July 8, 2026, but developers get no API access.Voice AI
Pipecat vs LiveKit Agents: The Trade-offs That Lock You InOriginalTwo open-source frameworks dominate production voice agents. They solve different stack layers and picking wrong means a rewrite. Here is how to decide.Voice AI
Grok Voice Gets 21 New Voices. The Price Is the PointOriginalxAI added 21 multilingual voices, voice cloning from one minute of audio, and a no-code agent builder to Grok Voice at $0.05 per minute of audio. Here's what that means for the voice AI market.Voice AI
An Open Source TTS Model That Edits Words After RecordingOriginalViiTorVoice-NAR is an open-source TTS model that can replace individual words inside finished audio without regenerating the surrounding content. It also clones voices without needing a transcript.Voice AI
Voice AI mid-2026: the trends I am watchingOriginalQuality convergence, pricing pressure, open-weight models, and the regulatory shift. The voice AI landscape halfway through 2026.Voice AI
5 voice AI stories that shaped the start of JulyOriginalOpenAI shipped voice reasoning, Anthropic models returned, ElevenLabs hit $22B, Deepgram went multilingual, and Google taught Gemini to use a computer.Voice AI
Someone Asked ChatGPT to Scream. It Did.OriginalA viral TikTok showed ChatGPT's Advanced Voice Mode screaming on command. Two screeches, one awkward silence, and a lot of questions about what we just watched.AI
What Voice Agent Pricing Reveals About the PlatformOriginalVoice agent platform pricing tells you more than cost. The pricing model reveals which layer of the stack each platform owns and what trade-offs you inherit.Voice AI
Notes from my first week at SpeechifyOriginalFirst week as Speechify's Head of DevRel: a TTS cookbook, two developer guides, and watching AI shift to government-gated releases.Developer Experience
Voice cloning goes open source, voice agents go enterpriseOriginalNetEase open-sourced voice cloning from 3 seconds of audio. ElevenLabs partnered with IBM and added SynthID. Coval raised $28M. UK laws are unfit.Voice AI
My Voice Was Cloned Before Lunch on Day OneOriginalDay one at Speechify and my voice was cloned before I finished onboarding. Hearing yourself through a TTS engine is a rite of passage I was not ready for.Voice AI
Voice AI's Real Competition Shifted From Models to PlatformsOriginalTTS model quality converged by June 2026. The real competitive moat shifted to developer experience, platform integration, and compliance tooling.Voice AI
TTS Latency: How Time to First Audio Actually WorksOriginalA deep dive into Time to First Audio, the metric that defines voice agent responsiveness, and what happens in the latency pipeline from text to speech.Voice AI
Six stories shaping voice AI in mid-June 2026OriginalMicrosoft MAI-Voice-2, Google live translation, DeepL bought Mixhalo, and open-weight TTS models kept shrinking. Six stories from a busy month in voice AI.Voice AI
When speech-to-text hears something else entirelyOriginalThe funniest and weirdest STT transcriptions from real Deepgram API usage. Some are bugs, some are features, and every single one made me laugh.Developer Experience
Voice data residency decides where your agent runsOriginalVoice data residency decides where speech APIs process audio. Deepgram Australia went live June 17, 2026, shaping how teams pick STT and TTS providers.Voice AI
Streaming TTS: Rethinking the Voice Audio PipelineOriginalStreaming TTS changes what voice applications expect from audio APIs. Time-to-first-byte drops, complexity moves into buffer management and chunk boundaries, and knowing when streaming isn't the answer matters just as much.Voice AI
When Your Drive-Thru AI Can't Understand an AccentOriginalMcDonald's pulled its AI drive-thru pilot in 2024 because it couldn't handle regional accents. By mid-2026, the industry was still figuring out why.Voice AI
Three STT Strategies, One MarketOriginalBy June 2026 every major STT provider hit acceptable accuracy. The real competition moved to multilingual support, turn detection and platform integration.Voice AI
Inside the Voice Agent Pipeline: STT, LLM, and Streaming TTSOriginalHow STT transcribes audio, an LLM generates responses, and streaming TTS speaks them back. A technical breakdown of the real-time pipeline behind voice agents in 2026.Voice AI
TIL: Flux natively knows when a caller finishes speakingOriginalDeepgram Flux turn detection replaces VAD and silence timeouts with model-native EndOfTurn events. A single WebSocket config parameter simplifies voice agent turn-taking.Voice AI
Testing Voice AI Means Talking to Yourself in PublicOriginalBuilding voice AI means reading test sentences aloud in coffee shops, on trains, and in meetings you forgot to mute. It looks as ridiculous as it sounds.Voice AI
When SDKs Write Themselves: Voice API Code GenerationOriginalDeepgram switched from hand-rolled SDKs to spec-first generation with Fern. Here is what that looked like across five languages and how it changed shipping voice APIs.Developer Experience
Sunday roundup: a quiet publishing week in voice AIOriginalTwo posts from late May: multilingual voice agent costs and why contribution guides exclude new contributors. Plus what else was happening in AI that week.Voice AI
5 command-line tools for shipping voice agentsOriginalBuilding voice agents means living in a terminal. Here are five CLI tools I use every day for audio debugging, API testing, and latency measurement.Developer Experience
What multilingual voice agents cost: latency and complexityOriginalBuilding a voice agent that handles ten languages without falling over is harder than it sounds. Here's what the architecture actually costs.Voice AI
Sunday roundup: API design, I/O, AnthropicOriginalOne post this week about voice API design. Google I/O and Anthropic's London event reshaped the AI landscape. Here is the roundup.Developer Experience
5 API design decisions that shape voice AI dev experienceOriginalError payloads, streaming edge cases, and latency limits all shape how developers interact with voice APIs. Here are five patterns I have seen matter most.Developer Experience
Testing 10 languages with one macOS commandOriginalThe macOS say command generates speech in ten languages. I used it to test a multilingual STT model without installing any audio tooling.TIL
What nobody tells you about audio in speech-to-textOriginalProduction speech-to-text needs preprocessing sample rate, encoding, and chunk sizes. The docs skip these. Years debugging production STT taught me what matters.Developer Experience
Voice AI APIs Converged Into Single-Call PlatformsOriginalThe week of May 11, 2026, three separate announcements pushed voice AI from multi-service pipelines toward unified single-call APIs. Here is what changed and why it matters for developers.Voice AI
Sunday roundup: four posts on audio debugging and work cultureOriginalFour posts from the week of May 11: TIL on afinfo, speaker diarization deep dive, Slack culture in remote teams, and ffmpeg for voice AI debugging. Plus what else was happening.Developer Experience
5 SDK anti-patterns I keep fixing in voice AIOriginalMaintaining SDKs across five languages taught me the same mistakes appear every time. Here are the five patterns I'd redesign first, and why they matter for voice AI.Developer Experience
ffmpeg taught me more about voice AI than the docs didOriginalThe first time I debugged a voice AI integration, ffmpeg saved me. It is still the most useful tool in my kit, and it is not even designed for voice.Developer Experience
Why Speaker Diarization Is the Hardest Problem in Voice AIOriginalSpeaker diarization figures out who spoke when. It sounds simple. It is not. Here is why it breaks, and what it takes to get right.Voice AI
TIL: What afinfo Reveals About Your Audio FilesOriginalafinfo prints every audio file property macOS knows about: sample rate, channels, bit depth, duration, and codec. Essential for debugging STT pipeline issues.TIL
5 developer experience wins in voice AI toolingOriginalError messages, timeout behavior, and observability patterns separate great voice APIs from frustrating ones. Here are five patterns that matter.Developer Experience
Friday fun: inspecting audio files from my terminalOriginalsoxi shows audio file metadata like sample rate, duration, and channels from the command line. Inspect files before sending them to any voice API.TIL
What running a developer Discord taught me about voice AIOriginalFour years in a voice AI developer community showed me the same problems again and again. Audio format issues, silent failures, and the questions nobody puts in the docs.Developer Experience
How voice AI SDKs handle things REST clients never have toOriginalBuilding SDKs for streaming voice APIs means managing WebSocket state, audio buffers, reconnection, and backpressure. REST client patterns break immediately.Developer Experience
Sunday roundup: debugging habits and multilingual speechOriginalOne post this week on voice AI debugging. Around it: Flux went multilingual, AssemblyAI launched a Voice Agent API, Twilio updated Conversation Relay, and the daily cadence began.Voice AI
5 habits that reduce voice AI debuggingOriginalMost voice AI debugging time goes to problems that follow a pattern. Here are five habits I built at Deepgram that catch those patterns before they become incidents.Engineering
Sunday roundup: starting fresh, April 26OriginalAnnouncing the start of daily publishing on lukeocodes.dev. Plus OpenAI workspace agents, Anthropic Mythos 5, and what was happening at Deepgram.Developer Experience
Friday fun: generating test tones for voice AI pipelinesOriginalsox synth generates test audio files from scratch at any sample rate and duration. No microphone needed for speech API development.TIL
Inside the Streaming Cascade Powering Voice AIOriginalReal-time voice agents chain three models over streaming connections. Here's how the cascade architecture works and where every millisecond goes.Voice AI
Sunday roundup: two posts, a fresh startOriginalTwo posts from the first weekend of daily writing on lukeocodes.dev. Terminal STT streaming, WebSocket audio patterns, and GPT-5.4-Cyber.Developer Experience
Five WebSocket patterns for real-time audio streamingOriginalStreaming audio WebSockets need reconnect logic, backpressure, heartbeats, and graceful shutdown. Five patterns covering the full connection lifecycle.Engineering
Friday fun: live speech to text from the terminalOriginalrec captures mic audio, websocat streams it to a WebSocket STT API, and the terminal shows the transcript in real time. No GUI needed.TIL