Luke Oliff.
Sunday roundup: six posts from a week in voice AIOriginalSix posts this week. EU watermarking, emotion prompts, open source voice agents, AI coding one year on, a raccoon heist game, and the AI industry roundup.Voice AI
TIL: Generating Audio Waveform Images With ffmpegOriginalffmpeg generates a waveform PNG from any audio file. Visualise silence gaps, volume levels, and speech patterns without a DAW.TIL
Voice Emotion Control Moves From SSML to PromptsOriginalVoice emotion control is moving to natural language prompts. Kakao's Kanana-o scores 94.50 on the Korean InstructTTSEval benchmark.Voice AI
EU AI Act Voice Watermarking: What TTS Builders Must KnowOriginalEU AI Act voice watermarking rules took effect August 2, 2026. What TTS providers and voice developers must do to stay compliant.Voice AI
WebSocket vs REST TTS APIs for Voice AgentsOriginalChoosing between WebSocket and REST for your streaming TTS API changes your voice agent's latency floor. Here is how to pick the right protocol.TTS
Smallest.ai Raises $13M to Split Voice Agents in TwoOriginalSmallest.ai closed a $13M Series A for Voice 4.0 and Hydra. Their bet is that fast real-time agents need small speech models, not big LLMs.Voice AI
Fish Audio's $52M Seed: Open Weights Got Them HereOriginalFish Audio raised a $52M seed on July 28, 2026 with $21M ARR and 8M users. Its new S2.1 Pro model is closed, API-only. What that shift means for TTS.TTS
What a Week at SpeechifyOriginalSimba went multilingual, streaming got timestamps, voice agents learned to switch languages mid-call. Qwen took the crown. MCP broke everything.Voice AI
MAI-Voice-Flash, Opus 5, and a Red LineOriginalMicrosoft MAI-Voice-2-Flash enters TTS at $15/M. Anthropic ships Opus 5 at half Fable's cost. OpenAI faces red-line questions after Hugging Face.Voice AI
TTS Quality Has No Single Number YetOriginalTTS quality has no single metric. MOS scores cluster, Elo only tests short English clips, and TTS WER measures a different thing than STT. Here is how to actually evaluate.TTS
TIL: Test TTS Voices With CurlOriginalOne curl command to compare TTS voice output side by side before you write any integration code. Handy for prototyping or auditing voice quality.TIL
GPT-Live Failed Its First Viral Test in 12 SecondsOriginalOpenAI's GPT-Live launched July 8 and the internet found its weak spot in hours. TikToker Husk broke it with a spelling test, and the full-duplex interruptions are already a meme.Voice AI
Grok Voice Gets 21 New Voices. The Price Is the PointOriginalxAI added 21 multilingual voices, voice cloning from one minute of audio, and a no-code agent builder to Grok Voice at $0.05 per minute of audio. Here's what that means for the voice AI market.Voice AI
An Open Source TTS Model That Edits Words After RecordingOriginalViiTorVoice-NAR is an open-source TTS model that can replace individual words inside finished audio without regenerating the surrounding content. It also clones voices without needing a transcript.Voice AI
TIL: Playing Audio From the Terminal With SoX PlayOriginalplay (from SoX) plays any audio file through your speakers with one command. No media player needed when you are already in the terminal.TIL
Pricing wars are good for developersOriginalThree TTS providers changed pricing in the same week. The trend is clear and it benefits everyone building with voice AI.Opinion
What surprised me about TTS API design after years of STTOriginalAfter years of building with speech-to-text APIs, switching to text-to-speech revealed design patterns I had never thought about.API Design
5 voice AI stories that shaped the start of JulyOriginalOpenAI shipped voice reasoning, Anthropic models returned, ElevenLabs hit $22B, Deepgram went multilingual, and Google taught Gemini to use a computer.Voice AI
Qwen-Audio-3.0-TTS Flash Comes for Real-Time VoiceOriginalAlibaba's Qwen-Audio-3.0-TTS Flash targets the real-time TTS market on price and latency. What it means for voice developers and the API field.AI
TIL: Clone a voice in one API call with SpeechifyOriginalThe Speechify API creates a cloned voice from a 10-second audio sample with a single POST, returning a voice ID that works on any speech endpoint you already use.TIL
Voice cloning goes open source, voice agents go enterpriseOriginalNetEase open-sourced voice cloning from 3 seconds of audio. ElevenLabs partnered with IBM and added SynthID. Coval raised $28M. UK laws are unfit.Voice AI
My Voice Was Cloned Before Lunch on Day OneOriginalDay one at Speechify and my voice was cloned before I finished onboarding. Hearing yourself through a TTS engine is a rite of passage I was not ready for.Voice AI
Voice AI's Real Competition Shifted From Models to PlatformsOriginalTTS model quality converged by June 2026. The real competitive moat shifted to developer experience, platform integration, and compliance tooling.Voice AI
TTS Latency: How Time to First Audio Actually WorksOriginalA deep dive into Time to First Audio, the metric that defines voice agent responsiveness, and what happens in the latency pipeline from text to speech.Voice AI
TIL: Read speech marks from the Speechify API responseOriginalEvery Speechify TTS response includes word-level timing data alongside the audio. Here is how to read speech marks and what you can do with them.TIL
Six stories shaping voice AI in mid-June 2026OriginalMicrosoft MAI-Voice-2, Google live translation, DeepL bought Mixhalo, and open-weight TTS models kept shrinking. Six stories from a busy month in voice AI.Voice AI
Voice data residency decides where your agent runsOriginalVoice data residency decides where speech APIs process audio. Deepgram Australia went live June 17, 2026, shaping how teams pick STT and TTS providers.Voice AI
Streaming TTS: Rethinking the Voice Audio PipelineOriginalStreaming TTS changes what voice applications expect from audio APIs. Time-to-first-byte drops, complexity moves into buffer management and chunk boundaries, and knowing when streaming isn't the answer matters just as much.Voice AI
Sunday roundup: four posts on audio debugging and work cultureOriginalFour posts from the week of May 11: TIL on afinfo, speaker diarization deep dive, Slack culture in remote teams, and ffmpeg for voice AI debugging. Plus what else was happening.Developer Experience
Sunday roundup: debugging habits and multilingual speechOriginalOne post this week on voice AI debugging. Around it: Flux went multilingual, AssemblyAI launched a Voice Agent API, Twilio updated Conversation Relay, and the daily cadence began.Voice AI