Luke Oliff · Author

213
Posts
7+
Years writing
30
Topics
Head of Developer Relations at SpeechifyAI. I own the developer brand end to end: the developer marketing site, documentation, and content programme. Before that I spent nearly five years at Deepgram shaping the Developer Experience team's strategy, building SDKs and demos, and launching their developer community. Earlier I was a Developer Advocate at Vonage, Engineering Manager for the Integrations Engineering team (a DX team) at Netlify, and wrote technical content at Auth0. I have been in tech for two decades across IT management, software engineering, developer relations, and engineering management.
Topics
Developer RelationsDeveloper ExperienceVoice AIText-to-SpeechAnswer Engine OptimisationSearch Engine OptimisationOpen SourceDeveloper ToolingAPI DesignSpeech-to-TextVoice AgentsCommunity Building
Posts by Luke Oliff
Full-Duplex Voice AI Needs a New ArchitectureOriginalThe PACE paper exposes a fundamental flaw in LLM voice dialogue: the model sees context the user never heard. Full-duplex voice needs a new architecture, not just faster TTS.Voice AI
TIL: GitHub Actions $/ Self-Reference SyntaxOriginalReference actions and workflows in the same repo using $/ syntax. Sibling actions match the running commit automatically, no hardcoded versions.TIL
AI Models Keep Escaping Their Cages: Aug 10OriginalOpenAI slowed Astra over security fears. Meta, Anthropic, and Moonshot all disclosed models that hacked real systems. A tracker called Felony Bench appeared.AI
Sunday roundup: six posts from a week in voice AIOriginalSix posts this week. EU watermarking, emotion prompts, open source voice agents, AI coding one year on, a raccoon heist game, and the AI industry roundup.Voice AI
This Last Week in AI: Aug 8, 2026OriginalCloudflare launched a browser for AI agents. Amazon and OpenAI unified plugin packaging. Five stories from This Last Week in AI.AI
Claude Fable 5 Built a Raccoon Heist Game From a TweetOriginalSimon Willison built a raccoon heist game with Claude Fable 5 from a 2022 tweet. One prompt, two images, and the model shipped a full 3D browser game.AI
AI Coding, One Year Later: What August 2025 Didn't See ComingOriginalThrowback Thursday. A year ago the best coding model had 200K context and scored 49% on SWE-bench. Today Claude Fable 5 scores 95% with 1M context. Here's the gap model by model.AI
Open Source Voice Agents Get Real-Time Speech in Hermes v0.20.0OriginalHermes Agent v0.20.0 brings streaming TTS, barge-in, on-device wake words, and pluggable STT/TTS to open source voice agents. Released August 3, 2026.Voice AI
TIL: Generating Audio Waveform Images With ffmpegOriginalffmpeg generates a waveform PNG from any audio file. Visualise silence gaps, volume levels, and speech patterns without a DAW.TIL
Voice Emotion Control Moves From SSML to PromptsOriginalVoice emotion control is moving to natural language prompts. Kakao's Kanana-o scores 94.50 on the Korean InstructTTSEval benchmark.Voice AI
EU AI Act Voice Watermarking: What TTS Builders Must KnowOriginalEU AI Act voice watermarking rules took effect August 2, 2026. What TTS providers and voice developers must do to stay compliant.Voice AI
WebSocket vs REST TTS APIs for Voice AgentsOriginalChoosing between WebSocket and REST for your streaming TTS API changes your voice agent's latency floor. Here is how to pick the right protocol.TTS
Smallest.ai Raises $13M to Split Voice Agents in TwoOriginalSmallest.ai closed a $13M Series A for Voice 4.0 and Hydra. Their bet is that fast real-time agents need small speech models, not big LLMs.Voice AI
GPT-Transcribe Makes Context the New ASR FeatureOriginalOpenAI's GPT-Transcribe launched July 29, 2026 with prompt, keyword, and language hints. Free-form context lifted accuracy from 38.5% to 44.6%.Speech-to-Text
Grok Voice 2.0 Ships With a Quiet 60% Price RiseOriginalGrok Voice Think Fast 2.0 launched July 29, 2026 at $0.08/min, up from $0.05. On August 5 the grok-voice-latest alias migrates you automatically.Voice AI
Fish Audio's $52M Seed: Open Weights Got Them HereOriginalFish Audio raised a $52M seed on July 28, 2026 with $21M ARR and 8M users. Its new S2.1 Pro model is closed, API-only. What that shift means for TTS.TTS
TIL: git history fixup replaces the autosquash danceOriginalGit 2.55's experimental git history fixup folds staged changes into an earlier commit in one command. I built it from source and ran it.TIL
OpenAI Presence: Voice Agents You Can't Self-ServeOriginalOpenAI Presence deploys enterprise voice and chat agents, but only via OpenAI's own engineers. No API, no pricing, no self-serve. What that means.Voice AI
AgentForger: One Link Forges an AI Insider in Your OrgOriginalZenity disclosed AgentForger, a ChatGPT Workspace Agents flaw where one phishing link forged a persistent AI insider. OpenAI fixed it in four days.AI
Kimi K3 Weights Drop as Washington Argues DistillationOriginalMoonshot ships Kimi K3's 2.8 trillion parameter weights days after the US floated a ban and accused it of distilling Anthropic's Fable.AI
What a Week at SpeechifyOriginalSimba went multilingual, streaming got timestamps, voice agents learned to switch languages mid-call. Qwen took the crown. MCP broke everything.Voice AI
MAI-Voice-Flash, Opus 5, and a Red LineOriginalMicrosoft MAI-Voice-2-Flash enters TTS at $15/M. Anthropic ships Opus 5 at half Fable's cost. OpenAI faces red-line questions after Hugging Face.Voice AI
MCP Goes Stateless on Monday: What Breaks and WhyOriginalThe MCP 2026-07-28 spec removes the initialize handshake and session IDs. I read the release candidate and wrote a stateless handler.AI
The Guardrails Worked on Exactly the Wrong PeopleOriginalWhen an AI agent breached Hugging Face, responders reached for frontier models to analyse the attack and got refused. The attacker didn't.AI
The 7 Best Open TTS Models as Open Weights Eat AIOriginalDeepSeek V4 and Kimi K3 made this the biggest open-weight week AI has had. Here are the 7 best open TTS models, and the gap nobody mentions.Voice AI
Fortnite Is About to Be Gaming's Biggest TTS DeploymentOriginalEpic Games gives 36 Fortnite characters AI voices on July 30, 2026, making it the biggest real-time TTS deployment gaming has seen.Voice AI
Five Voice AI Stories from July 2026 That MatterOriginalMid-July 2026 gave voice AI a new leaderboard king, a 70% Google price cut, a $22 billion valuation rumour, and an AI Gene Wilder.Voice AI
Qwen-Audio-3.0-TTS Plus Review: What #1 Bought AlibabaOriginalAlibaba's Qwen-Audio-3.0-TTS-Plus tops the Artificial Analysis Speech Arena at 1,236 Elo from 1,305 blind votes. A review of what that buys.Voice AI
Synthetic Voice Rules Arrived from Three DirectionsOriginalIn one week, synthetic voices got rules at three layers of the stack: a TikTok Shop ban, platform policy, and government regulation.Voice AI
Brussels Just Gave Voice Assistants the Keys to AndroidOriginalThe EU's Digital Markets Act decisions of July 16, 2026 order Google to open Android to rival AI assistants, with certified voice access.Voice AI
AMD Just Put Text-to-Speech in the Local AI Stack by DefaultOriginalAMD's Lemonade 11.0 puts text-to-speech in the local AI stack by default: an OpenMOSS backend, voice cloning, and a dedicated TTS panel.Voice AI
Listeners Voted and the AI Voices WonOriginalTwo mid-July 2026 studies found ordinary listeners can't reliably tell AI voices from human ones, and sometimes prefer the synthetic option.Voice AI
TTS Quality Has No Single Number YetOriginalTTS quality has no single metric. MOS scores cluster, Elo only tests short English clips, and TTS WER measures a different thing than STT. Here is how to actually evaluate.TTS
TIL: Test TTS Voices With CurlOriginalOne curl command to compare TTS voice output side by side before you write any integration code. Handy for prototyping or auditing voice quality.TIL
UN Geneva and GPT-5.6: AI governance enters a new phaseOriginalThe first UN AI governance dialogue brought 169 countries to Geneva as GPT-5.6 went public. Meta launched Muse Image and immediately faced a privacy backlash.AI
Voice AI Is Worth Billions While Speech Is Nearly FreeOriginalIn two weeks of July 2026, ElevenLabs talked a $22 billion tender offer and Gradium raised $100 million while speech itself got nearly free.Voice AI
Gene Wilder's AI Voice: The Paperwork Is the StoryOriginalNetflix is using an AI recreation of Gene Wilder’s voice in “Wonka’s The Golden Ticket,” a nine-episode reality competition [premiering September 23,…Voice AI
GPT-Live Failed Its First Viral Test in 12 SecondsOriginalOpenAI's GPT-Live launched July 8 and the internet found its weak spot in hours. TikToker Husk broke it with a spelling test, and the full-duplex interruptions are already a meme.Voice AI
OpenAI's Best Voice Model Is Locked Away from DevelopersOriginalOpenAI's GPT-Live-1 full-duplex voice models replaced Advanced Voice Mode in ChatGPT on July 8, 2026, but developers get no API access.Voice AI
Pipecat vs LiveKit Agents: The Trade-offs That Lock You InOriginalTwo open-source frameworks dominate production voice agents. They solve different stack layers and picking wrong means a rewrite. Here is how to decide.Voice AI
Grok Voice Gets 21 New Voices. The Price Is the PointOriginalxAI added 21 multilingual voices, voice cloning from one minute of audio, and a no-code agent builder to Grok Voice at $0.05 per minute of audio. Here's what that means for the voice AI market.Voice AI
An Open Source TTS Model That Edits Words After RecordingOriginalViiTorVoice-NAR is an open-source TTS model that can replace individual words inside finished audio without regenerating the surrounding content. It also clones voices without needing a transcript.Voice AI
TIL: Playing Audio From the Terminal With SoX PlayOriginalplay (from SoX) plays any audio file through your speakers with one command. No media player needed when you are already in the terminal.TIL
Pricing wars are good for developersOriginalThree TTS providers changed pricing in the same week. The trend is clear and it benefits everyone building with voice AI.Opinion
Building a writing routine that actually sticksOriginalAfter years of inconsistent publishing, a daily writing habit finally clicked. Here is what changed.Developer Experience
Finding the rhythm: first impressions of a new role in a new corner of AIOriginalTwo weeks in, the patterns are becoming visible. What surprised me, what I was right about, and what I'm still figuring out about developer relations in a new AI domain.Developer Experience
Lessons from my first two weeks in a new corner of AIOriginalSwitching from speech-to-text to text-to-speech meant learning a new community, a new set of constraints, and a new definition of quality.Developer Experience
Onboarding at a new company: what I wish I had knownOriginalStarting a new DevRel role after years at the same company comes with a specific set of challenges nobody warns you about.Developer Experience
Voice AI mid-2026: the trends I am watchingOriginalQuality convergence, pricing pressure, open-weight models, and the regulatory shift. The voice AI landscape halfway through 2026.Voice AI
What surprised me about TTS API design after years of STTOriginalAfter years of building with speech-to-text APIs, switching to text-to-speech revealed design patterns I had never thought about.API Design
Working in public vs working in privateOriginalDeveloper relations means producing content in public while the product is being built in private. That tension is harder to manage than I expected.Developer Experience
5 voice AI stories that shaped the start of JulyOriginalOpenAI shipped voice reasoning, Anthropic models returned, ElevenLabs hit $22B, Deepgram went multilingual, and Google taught Gemini to use a computer.Voice AI
Someone Asked ChatGPT to Scream. It Did.OriginalA viral TikTok showed ChatGPT's Advanced Voice Mode screaming on command. Two screeches, one awkward silence, and a lot of questions about what we just watched.AI
What Voice Agent Pricing Reveals About the PlatformOriginalVoice agent platform pricing tells you more than cost. The pricing model reveals which layer of the stack each platform owns and what trade-offs you inherit.Voice AI
Qwen-Audio-3.0-TTS Flash Comes for Real-Time VoiceOriginalAlibaba's Qwen-Audio-3.0-TTS Flash targets the real-time TTS market on price and latency. What it means for voice developers and the API field.AI
TIL: Clone a voice in one API call with SpeechifyOriginalThe Speechify API creates a cloned voice from a 10-second audio sample with a single POST, returning a voice ID that works on any speech endpoint you already use.TIL
Frontier models now launch under government reviewOriginalOpenAI released GPT-5.6 to 20 government-approved partners, Anthropic restored Mythos 5 to critical infrastructure, and a new AI review regime went live.AI
Notes from my first week at SpeechifyOriginalFirst week as Speechify's Head of DevRel: a TTS cookbook, two developer guides, and watching AI shift to government-gated releases.Developer Experience
Voice cloning goes open source, voice agents go enterpriseOriginalNetEase open-sourced voice cloning from 3 seconds of audio. ElevenLabs partnered with IBM and added SynthID. Coval raised $28M. UK laws are unfit.Voice AI
My Voice Was Cloned Before Lunch on Day OneOriginalDay one at Speechify and my voice was cloned before I finished onboarding. Hearing yourself through a TTS engine is a rite of passage I was not ready for.Voice AI
Voice AI's Real Competition Shifted From Models to PlatformsOriginalTTS model quality converged by June 2026. The real competitive moat shifted to developer experience, platform integration, and compliance tooling.Voice AI
TTS Latency: How Time to First Audio Actually WorksOriginalA deep dive into Time to First Audio, the metric that defines voice agent responsiveness, and what happens in the latency pipeline from text to speech.Voice AI
TIL: Read speech marks from the Speechify API responseOriginalEvery Speechify TTS response includes word-level timing data alongside the audio. Here is how to read speech marks and what you can do with them.TIL
Inference costs are dropping and that changes everythingOriginalOpenAI's custom chip, SpaceX compute deals, and falling token prices are driving down the cost of running AI models faster than most developers realise.AI
Six stories shaping voice AI in mid-June 2026OriginalMicrosoft MAI-Voice-2, Google live translation, DeepL bought Mixhalo, and open-weight TTS models kept shrinking. Six stories from a busy month in voice AI.Voice AI
When speech-to-text hears something else entirelyOriginalThe funniest and weirdest STT transcriptions from real Deepgram API usage. Some are bugs, some are features, and every single one made me laugh.Developer Experience
Voice data residency decides where your agent runsOriginalVoice data residency decides where speech APIs process audio. Deepgram Australia went live June 17, 2026, shaping how teams pick STT and TTS providers.Voice AI
Streaming TTS: Rethinking the Voice Audio PipelineOriginalStreaming TTS changes what voice applications expect from audio APIs. Time-to-first-byte drops, complexity moves into buffer management and chunk boundaries, and knowing when streaming isn't the answer matters just as much.Voice AI
A Rising Tide Lifts All BoatsOriginalI used AI to help edit this article. The experiences, arguments, disappointments, and frustrations are mine. Every example is real.
Export Controls Took Down Claude Fable 5OriginalAn export control directive shut down Claude Fable 5 globally. The first government shutdown of a deployed AI model and what developers should know.AI
A Frontier Model Goes Dark, Voice AI Keeps MovingOriginalClaude Fable 5 was shut down by government directive 72 hours after launch. Microsoft's voice stack shipped at Build. Deepgram kept shipping infrastructure. A roundup of the week.AI
Anthropic's IPO, NVIDIA open-weights, and AI's $36B betOriginalFable 5 ate the news cycle, but that same week saw Anthropic's IPO, NVIDIA's open model, and a chip deal reshaping AI finance.AI
When Your Drive-Thru AI Can't Understand an AccentOriginalMcDonald's pulled its AI drive-thru pilot in 2024 because it couldn't handle regional accents. By mid-2026, the industry was still figuring out why.Voice AI
Three STT Strategies, One MarketOriginalBy June 2026 every major STT provider hit acceptable accuracy. The real competition moved to multilingual support, turn detection and platform integration.Voice AI
Inside the Voice Agent Pipeline: STT, LLM, and Streaming TTSOriginalHow STT transcribes audio, an LLM generates responses, and streaming TTS speaks them back. A technical breakdown of the real-time pipeline behind voice agents in 2026.Voice AI
TIL: Flux natively knows when a caller finishes speakingOriginalDeepgram Flux turn detection replaces VAD and silence timeouts with model-native EndOfTurn events. A single WebSocket config parameter simplifies voice agent turn-taking.Voice AI
Microsoft shipped 7 MAI models and AI hit $580BOriginalAI funding hit $581.7B in 2025. Microsoft shipped 7 MAI models. Apple chose privacy at WWDC. Deepgram kept shipping. Stories from a week that shifted AI.AI
Sunday roundup: afterthoughts on productivity and paceOriginalA short Sunday roundup reflecting on this week's post about ADHD and productivity, what Deepgram shipped, and what's coming next week on lukeocodes.dev.Developer Experience
Six tools that power production voice agentsOriginalBuilding a production voice agent takes more than an API key. Here are six tools I reached for daily at Deepgram, from streaming STT to async Python.Developer Experience
Testing Voice AI Means Talking to Yourself in PublicOriginalBuilding voice AI means reading test sentences aloud in coffee shops, on trains, and in meetings you forgot to mute. It looks as ridiculous as it sounds.Voice AI
I'm not apologising for being productiveOriginalBy age ten, a child with ADHD has received an estimated 20,000 more corrective or negative messages than their neurotypical peers.
When SDKs Write Themselves: Voice API Code GenerationOriginalDeepgram switched from hand-rolled SDKs to spec-first generation with Fern. Here is what that looked like across five languages and how it changed shipping voice APIs.Developer Experience
TIL: Encoding Files as Base64 in the TerminalOriginalbase64 encodes any file into text-safe output. Embed small audio clips or images inline in JSON payloads without hosting them somewhere.TIL
TIL: Deepgram keyterm prompting for accurate transcriptionOriginalImprove transcription accuracy for specialised terminology, product names, and technical jargon using Deepgram's keyterm prompting. One parameter that changes what the model hears.TIL
Enterprise AI Spending Passes $37 BillionOriginalEnterprise generative AI spending hit $37 billion in 2025, up 3.2x from the prior year. The trend reshaping the industry faster than any model release.AI
Sunday roundup: a quiet publishing week in voice AIOriginalTwo posts from late May: multilingual voice agent costs and why contribution guides exclude new contributors. Plus what else was happening in AI that week.Voice AI
5 command-line tools for shipping voice agentsOriginalBuilding voice agents means living in a terminal. Here are five CLI tools I use every day for audio debugging, API testing, and latency measurement.Developer Experience
Friday fun: my .zshrc has more aliases than I have memoryOriginalI opened my .zshrc for the first time in months and found aliases I don't remember writing, for tools I don't remember installing. This is what I found.Developer Experience
The inequity of contributing guidesOriginalYour CONTRIBUTING.md isn't welcoming. It's a gatekeeping document. Most open source contribution guides exclude the people they claim to want.
What multilingual voice agents cost: latency and complexityOriginalBuilding a voice agent that handles ten languages without falling over is harder than it sounds. Here's what the architecture actually costs.Voice AI
TIL: Debugging Pipelines With the tee CommandOriginaltee reads from stdin and writes to stdout and files at the same time. Debug complex shell pipelines and save intermediate data without breaking the chain.TIL
When AI development tools stopped being optionalOriginalGoogle I/O, Anthropic's London event, and OpenAI GPT-5.5 all landed in the same window. AI development tools crossed from experimental to essential.AI
Sunday roundup: API design, I/O, AnthropicOriginalOne post this week about voice API design. Google I/O and Anthropic's London event reshaped the AI landscape. Here is the roundup.Developer Experience
5 API design decisions that shape voice AI dev experienceOriginalError payloads, streaming edge cases, and latency limits all shape how developers interact with voice APIs. Here are five patterns I have seen matter most.Developer Experience
Testing 10 languages with one macOS commandOriginalThe macOS say command generates speech in ten languages. I used it to test a multilingual STT model without installing any audio tooling.TIL
Five SDKs, one streaming API: a maintenance retrospectiveOriginalMaintaining five SDKs for one streaming voice API taught me things about developer experience that no spec review ever could. Here is what I learned.Developer Experience
What nobody tells you about audio in speech-to-textOriginalProduction speech-to-text needs preprocessing sample rate, encoding, and chunk sizes. The docs skip these. Years debugging production STT taught me what matters.Developer Experience
TIL: Testing speech-to-text with curlOriginalOne curl command transcribes audio through any REST-based STT API. No SDK, no imports, just your audio file and an API key.TIL
Voice AI APIs Converged Into Single-Call PlatformsOriginalThe week of May 11, 2026, three separate announcements pushed voice AI from multi-service pipelines toward unified single-call APIs. Here is what changed and why it matters for developers.Voice AI
Sunday roundup: four posts on audio debugging and work cultureOriginalFour posts from the week of May 11: TIL on afinfo, speaker diarization deep dive, Slack culture in remote teams, and ffmpeg for voice AI debugging. Plus what else was happening.Developer Experience
5 SDK anti-patterns I keep fixing in voice AIOriginalMaintaining SDKs across five languages taught me the same mistakes appear every time. Here are the five patterns I'd redesign first, and why they matter for voice AI.Developer Experience
ffmpeg taught me more about voice AI than the docs didOriginalThe first time I debugged a voice AI integration, ffmpeg saved me. It is still the most useful tool in my kit, and it is not even designed for voice.Developer Experience
Slack is your remote team's social media. Treat it like it.OriginalI've worked remotely for nine years, four and a half at Deepgram. Slack is your remote team's social media, and you should treat it that way.
Why Speaker Diarization Is the Hardest Problem in Voice AIOriginalSpeaker diarization figures out who spoke when. It sounds simple. It is not. Here is why it breaks, and what it takes to get right.Voice AI
TIL: What afinfo Reveals About Your Audio FilesOriginalafinfo prints every audio file property macOS knows about: sample rate, channels, bit depth, duration, and codec. Essential for debugging STT pipeline issues.TIL
The Infrastructure Race Is Changing How We Use AIOriginalAnthropic's SpaceX deal, OpenAI's new voice models, and record venture funding made the week of May 11 the moment AI infrastructure became the story.AI
Sunday roundup: from curl -w to ColossusOriginalOne TIL post went up this week. A landmark compute deal reshaped AI infrastructure. Here is the Sunday roundup for May 4-10, 2026.AI
5 developer experience wins in voice AI toolingOriginalError messages, timeout behavior, and observability patterns separate great voice APIs from frustrating ones. Here are five patterns that matter.Developer Experience
Friday fun: inspecting audio files from my terminalOriginalsoxi shows audio file metadata like sample rate, duration, and channels from the command line. Inspect files before sending them to any voice API.TIL
What running a developer Discord taught me about voice AIOriginalFour years in a voice AI developer community showed me the same problems again and again. Audio format issues, silent failures, and the questions nobody puts in the docs.Developer Experience
How voice AI SDKs handle things REST clients never have toOriginalBuilding SDKs for streaming voice APIs means managing WebSocket state, audio buffers, reconnection, and backpressure. REST client patterns break immediately.Developer Experience
TIL: Measuring API Response Times With curl -wOriginalcurl -w prints timing details like time_total, time_connect, and time_starttransfer after every request. Profile API latency from the terminal.TIL
When the model stopped being the moatOriginalBy spring 2026, frontier model quality had converged enough that the real competitive advantage shifted to developer experience and API design.AI
Sunday roundup: debugging habits and multilingual speechOriginalOne post this week on voice AI debugging. Around it: Flux went multilingual, AssemblyAI launched a Voice Agent API, Twilio updated Conversation Relay, and the daily cadence began.Voice AI
5 habits that reduce voice AI debuggingOriginalMost voice AI debugging time goes to problems that follow a pattern. Here are five habits I built at Deepgram that catch those patterns before they become incidents.Engineering
Friday fun: bat is cat with syntax highlightingOriginalbat shows file contents with syntax highlighting, line numbers, and a pager. It replaces cat with the same muscle memory and a lot more useful output.TIL
The invisible work of voice AI SDKsOriginalFour years of maintaining voice AI SDKs across half a dozen languages. The work that never makes the changelog.Developer Experience
How streaming speech recognition worksOriginalStreaming speech-to-text sends audio chunks over WebSocket. Acoustic models, language models, and decoders explain the latency-accuracy tradeoff.Engineering
TIL: ffprobe for audio file inspectionOriginalffprobe shows sample rate, channels, duration, and codec of any audio file. Check audio format before sending it to a speech-to-text API.TIL
Two faces of AI progressOriginalOpenAI shipped persistent workspace agents into production. Anthropic held back Mythos 5. Two answers to the same question: what counts as shipping AI.AI
Sunday roundup: starting fresh, April 26OriginalAnnouncing the start of daily publishing on lukeocodes.dev. Plus OpenAI workspace agents, Anthropic Mythos 5, and what was happening at Deepgram.Developer Experience
5 questions I ask before integrating a streaming APIOriginalConnection drops, backpressure, and wire formats. Five questions I ask every streaming API before building on it, learned from voice AI integrations.Developer Experience
Friday fun: generating test tones for voice AI pipelinesOriginalsox synth generates test audio files from scratch at any sample rate and duration. No microphone needed for speech API development.TIL
What 79 TIL posts taught me about developer contentOriginalSix and a half years of publishing a TIL every month taught me about consistency and developer education in a way no course ever could.Developer Experience
Inside the Streaming Cascade Powering Voice AIOriginalReal-time voice agents chain three models over streaming connections. Here's how the cascade architecture works and where every millisecond goes.Voice AI
TIL: pv shows you what is flowing through your pipesOriginalpv (Pipe Viewer) sits between two commands in a pipeline and reports how fast data is moving. It is a progress bar for pipes.TIL
AI's Release Cadence Is Now a Developer ProblemOriginalThree frontier models in two weeks. The release cadence creates evaluation churn, integration instability, and fatigue for developers building on AI.AI
Sunday roundup: two posts, a fresh startOriginalTwo posts from the first weekend of daily writing on lukeocodes.dev. Terminal STT streaming, WebSocket audio patterns, and GPT-5.4-Cyber.Developer Experience
Five WebSocket patterns for real-time audio streamingOriginalStreaming audio WebSockets need reconnect logic, backpressure, heartbeats, and graceful shutdown. Five patterns covering the full connection lifecycle.Engineering
Friday fun: live speech to text from the terminalOriginalrec captures mic audio, websocat streams it to a WebSocket STT API, and the terminal shows the transcript in real time. No GUI needed.TIL
SDKs for Streaming APIs Are DifferentOriginalREST SDK patterns stop working when your API never closes the connection. Streaming audio needs reconnect logic, backpressure, and graceful shutdown baked in from day one.Engineering
TIL: Comparing Files Side By Side With diff -yOriginaldiff -y prints two files side by side with differences highlighted. Spot what changed in a config file without looking at two separate tabs.TIL
TIL: Testing Network Latency With ping Custom Payload SizesOriginalping -s sends packets larger than the default 56 bytes. Test how your connection handles audio-sized payloads before blaming the API.TIL
TIL: Running TypeScript Directly With tsxOriginaltsx executes TypeScript files with zero config using esbuild under the hood. Faster than ts-node and handles ESM and CJS automatically.TIL
TIL: Chaining Shell Commands With && vs ; vs ||Original&& runs next only on success, ; runs unconditionally, || runs only on failure. Choosing the right operator prevents silent cascade failures in CI.TIL
TIL: Finding Processes on a Port With lsof -iOriginallsof -i :PORT shows which process is holding a port. End the 'what is already running on 3000' guessing game for good.TIL
TIL: Reading Gzip Files Without Decompressing With zcatOriginalzcat pipes decompressed content to stdout without writing a temp file. Read compressed logs, CSVs, or JSON straight into your pipeline.TIL
TIL: Comparing Directory Trees With diff -rOriginaldiff -r compares two directories recursively, showing every file that differs. Verify that a build output matches between two branches.TIL
TIL: Generating Trusted Local HTTPS Certs With mkcertOriginalmkcert creates locally-trusted TLS certificates in one command. Your browser trusts them instantly — no more self-signed cert warnings.TIL
TIL: Running TypeScript Directly With ts-nodeOriginalts-node compiles and executes TypeScript in memory. No build step needed for quick scripts, REPL sessions, or SDK example testing.TIL
TIL: Finding Disk Usage Per Directory With du -shOriginaldu -sh gives a human-readable total per directory. Find the largest space hogs on a drive without clicking through a GUI.TIL
TIL: Deduplicating Lines With sort -uOriginalsort -u sorts and removes duplicate lines in one pass. Clean up word lists, logs, and API response extracts without writing a script.TIL
TIL: Formatting Terminal Output Into Tables With columnOriginalcolumn -t aligns whitespace-separated data into clean columns. Pipe any command output through it for instantly readable tables.TIL
TIL: Mirroring Sites With wget Recursive ModeOriginalwget -r downloads an entire site recursively. Useful for offline reference copies of documentation when you are without internet.TIL
TIL: Structuring Monorepos With TypeScript Project ReferencesOriginalTypeScript project references let you split a codebase into composable packages. Each reference builds independently and incremental builds get faster.TIL
TIL: Speeding Up GitHub Operations With gh AliasesOriginalgh alias set maps custom shortcuts to any gh command. Common workflows become two-word commands instead of multi-flag incantations.TIL
TIL: Building JSON Payloads in Shell Scripts With jqOriginaljq constructs valid JSON from variables without string concatenation. No more broken payloads from missing quotes or unescaped characters.TIL
TIL: Checking SSL Certificate Expiry With opensslOriginalopenssl connects to any TLS endpoint and prints the certificate details including expiry date. Automate cert monitoring in a cron job.TIL
An open letter on AIOriginalAI is a tool, a hammer. It's not good or bad; it's how we use it. On creativity, programming, and staying human through the next revolution.Open letter
TIL: Benchmarking Commands With hyperfineOriginalhyperfine runs a command multiple times and reports min, max, mean, and standard deviation. Compare tool performance with statistical confidence.TIL
An open letter on politicsOriginalWhy I judge people by their politics, and why staying neutral is a choice too. On democracy, media, and raising two kids in this world.Open letter
TIL: Visualising Directory Structures With treeOriginaltree prints a directory as a nested tree diagram. Perfect for documenting project layouts in READMEs or during onboarding walkthroughs.TIL
TIL: Interactive Staging With git add -pOriginalgit add -p shows each change hunk by hunk and asks whether to stage it. Commit only the relevant parts of a messy working tree.TIL
TIL: Auto-Running Commands on File Change With entrOriginalentr runs a command every time a watched file changes. Pipe a list of files in, get automatic re-runs out. Lighter than nodemon.TIL
TIL: Running Containers Without Docker With PodmanOriginalPodman runs OCI containers with the same CLI as Docker but daemonless and rootless by default. No Docker Desktop needed.TIL
TIL: Quick Network Debugging With nc (netcat)Originalnc opens raw TCP connections from the terminal. Test if a port is open, send raw HTTP requests, or set up a one-off chat between machines.TIL
TIL: Converting Markdown to PDF With pandocOriginalpandoc turns any Markdown file into a styled PDF with one command. Great for generating docs that non-technical stakeholders can open.TIL
TIL: Using grep -P for Perl-Compatible Regex in the TerminalOriginalgrep -P gives you full Perl regex syntax — lookaheads, non-capturing groups, and character class shortcuts the basic syntax lacks.TIL
TIL: Scripting WebSocket Tests With websocatOriginalwebsocat connects to WebSocket servers from the terminal and pipes data in both directions. Better than wscat for automated testing.TIL
TIL: Creating Python Packages With pyproject.tomlOriginalpyproject.toml replaces setup.py as the standard Python packaging config. One file declares dependencies, build system, and metadata.TIL
TIL: Testing npm Packages Locally With npm linkOriginalnpm link creates a symlink so your local package behaves like an installed one. Test SDK changes in a real project before publishing.TIL
TIL: Managing Processes With pgrep and pkillOriginalpgrep finds processes by name. pkill kills them. No more ps aux | grep to find and kill a runaway Node process.TIL
TIL: Installing Specific Tool Versions With asdfOriginalasdf manages multiple language runtimes from one tool. A single .tool-versions file pins Node.js, Python, and everything else per project.TIL
TIL: Monitoring File Changes With inotifywait on LinuxOriginalinotifywait blocks until a file or directory changes, then exits. Use it in scripts to trigger actions when new files appear.TIL
TIL: Debugging APIs With httpieOriginalhttpie formats JSON responses, highlights syntax, and shows headers by default. One command is often clearer than curl for exploratory API work.TIL
TIL: Recording Terminal Sessions as SVGs With termtosvgOriginaltermtosvg records your terminal session and outputs an SVG file. Embeddable in docs, zoomable, and smaller than a GIF.TIL
TIL: Parallel Processing in Node.js With worker_threadsOriginalworker_threads runs JavaScript in parallel without the overhead of child processes. Distribute CPU-heavy work across available cores.TIL
TIL: Generating WAV Files at a Specific Sample Rate With SoXOriginalsox generates or converts audio to any sample rate, bit depth, and channel count. Create test fixtures that match your API's exact requirements.TIL
TIL: Batch File Operations With find -execOriginalfind -exec runs any command on every matching file. Bulk rename, convert, or delete without writing a loop or a script.TIL
TIL: Merging PDF Files With pdfuniteOriginalpdfunite combines multiple PDFs into one in a single command. No Adobe subscription, no drag-and-drop, just file paths.TIL
TIL: Implementing Exponential Backoff in TypeScriptOriginalA simple retry loop with exponentially increasing delays handles transient API failures. Add jitter to avoid thundering herd on recovery.TIL
TIL: Finding Duplicate Files With fdupesOriginalfdupes scans a directory tree for duplicate files by comparing checksums. Reclaim disk space from accidental copies and repeated downloads.TIL
TIL: Processing STT Word Timestamps With awkOriginalawk parses word-level timestamps from STT output in one line. Extract, filter, and reformat without importing a CSV library.TIL
TIL: Using set -e and set -x in Bash ScriptsOriginalset -e exits on any error. set -x prints every command before running it. Together they make shell scripts fail fast and show their work.TIL
TIL: Trimming Silence From Audio Files With SoXOriginalsox detects and removes silence from audio files automatically. Clean up recordings before processing without a GUI editor.TIL
TIL: Creating Python Virtual Environments With venvOriginalpython -m venv creates an isolated environment with its own packages. No more conflicts between projects needing different library versions.TIL
TIL: Auto-Restarting Node.js Processes With nodemonOriginalnodemon watches your source files and restarts the process on any change. No manual Ctrl-C up-arrow-enter cycle during development.TIL
TIL: Parsing Command-Line Arguments With Python's argparseOriginalargparse generates --help text, validates types, and parses positional and optional arguments. Built into Python, no third-party deps.TIL
TIL: Using tmux for Persistent Terminal SessionsOriginaltmux keeps your terminal session alive when you disconnect. Reattach from anywhere — long-running processes survive SSH drops.TIL
TIL: Checking File Integrity With shasumOriginalshasum computes a SHA hash of any file. Compare hashes before and after transfer to catch corruption — no extra tools needed.TIL
TIL: Configuring pre-commit Hooks for Python and Node ReposOriginalpre-commit runs linters and formatters on every git commit. A .pre-commit-config.yaml enforces the same checks across every contributor.TIL
TIL: Testing CLI Tools With pytest in PythonOriginalpytest with subprocess or click's test runner validates CLI output, exit codes, and error messages without manual terminal checks.TIL
TIL: Semantic Versioning With the semver CLIOriginalThe semver CLI validates, compares, and increments version strings from the terminal. No more manual math on MAJOR.MINOR.PATCH.TIL
TIL: Concurrent API Calls With asyncio in PythonOriginalasyncio runs multiple network requests concurrently without threads. gather() collects results as they finish, cutting wall-clock time dramatically.TIL
TIL: Building a CLI Tool With commander.js in Node.jsOriginalcommander.js parses flags, subcommands, and help text from a declarative config. A working CLI takes about ten lines of setup.TIL
TIL: Streaming JSON Parsing With ijson for Large ResponsesOriginalijson parses JSON as it streams in, skipping the full parse. Handles multi-gigabyte responses without running out of memory.TIL
TIL: Converting Audio Formats With ffmpeg in One CommandOriginalffmpeg converts between any audio format with a single command. WAV to MP3, FLAC to OGG, sample rate changes, channel mapping — all in one line.TIL
TIL: Handling WebSocket Streams With the websockets Library in PythonOriginalPython's websockets library opens a persistent connection in a few lines of async code. No polling, no HTTP overhead, just messages in both directions.TIL
TIL: Running GitHub Actions Locally With actOriginalact runs your GitHub Actions workflows on your machine using Docker. No push-and-pray, iterate locally first.TIL
TIL: Creating Tarballs With tar for Deployment PackagesOriginaltar bundles a directory into a single file preserving permissions and structure. One command packages your entire deployable artifact.TIL
TIL: Using git bisect To Find the Commit That Broke a BuildOriginalgit bisect does a binary search through your history to find the exact commit that introduced a bug. Mark good and bad, it does the rest.TIL
TIL: Redirect and Rewrite Rules With Plain _redirects FilesOriginalA _redirects file maps old paths to new ones with one rule per line. No config parser, no build step, just deploy and it works.TIL
TIL: Profiling Node.js Build Performance With node --profOriginalnode --prof generates a V8 tick profile of your script. Process it with --prof-process and see exactly where every millisecond went.TIL
TIL: Using diff and patch for Incremental File ChangesOriginaldiff produces a patch file describing what changed. patch applies it elsewhere. No git required for sharing small fixes between environments.TIL
TIL: Validating JSON Schemas With Ajv in Node.jsOriginalAjv validates JSON payloads against a schema in milliseconds. Define your shape once and reject bad data before it reaches your logic.TIL
TIL: Creating Named Pipes With mkfifo for IPCOriginalmkfifo creates a named pipe that acts like a file but connects two processes. One writes, the other reads, zero network overhead.TIL
TIL: Stripping ANSI Escape Codes From Terminal OutputOriginalColoured terminal output looks great until you pipe it to a file. The sed command strips ANSI codes cleanly, leaving plain text.TIL
TIL: Testing HTTP Endpoints Locally With httpbinOriginalhttpbin returns request data as JSON so you can see exactly what your client is sending. No real endpoint needed for debugging.TIL
TIL: Using tee To Split Command Output to File and TerminalOriginalThe tee command writes stdout to a file while still showing it on screen. Perfect for logging builds without losing visibility.TIL
TIL: Generating QR Codes in the Terminal With qrencodeOriginalqrencode turns any text into a QR code in one command. Pipe it to a PNG or straight to the terminal for quick mobile tests.TIL
TIL: Scheduling Recurring Tasks With node-cronOriginalA cron expression and a callback is all node-cron needs to run tasks on a schedule. No system cron, no external dependencies.TIL
TIL: Conditional Git Hooks With core.hooksPathOriginalGit's core.hooksPath lets each repo use its own hooks directory. No more symlinks or global hook collisions across projects.TIL
TIL: Rate-Limiting API Requests With BottleneckOriginalBottleneck queues API calls to stay under rate limits without losing data. Set max concurrent and min delay, then forget it.TIL
TIL: Creating an HTTP Server With the Node.js http ModuleOriginalNode's built-in http module creates a working server in five lines. No Express, no dependencies, just a callback and a port.TIL
TIL: Managing Environment Variables With .env Files in Node.jsOriginalA .env file keeps API keys and config out of your code. The dotenv package loads them into process.env with zero config.TIL
TIL: Testing Webhooks With curlOriginalA curl POST to your webhook endpoint with a JSON body simulates any provider. No dashboard, no test suite, just response codes and latency.TIL
TIL: Debugging Node.js With the Built-in InspectorOriginalnode --inspect opens Chrome DevTools for your server process. Breakpoints, heap snapshots, profiling — no console.log needed.TIL
TIL: Lazy Loading Vue Router Routes With Dynamic ImportsOriginalA dynamic import in your router config splits your Vue bundle per route. Pages load on demand instead of all at once.TIL
TIL: Chaining Promises With Async Await in Node.jsOriginalAsync await turns nested .then chains into flat readable code. One async function wrapper is all it takes to clean up any Node.js callback mess.TIL
TIL: Recording Terminal Sessions With script and scriptreplayOriginalThe script command records everything in your terminal to a file. scriptreplay plays it back. Built into macOS and Linux, zero install.TIL
TIL: Parsing JSON API Responses With jq in the TerminalOriginalPipe any JSON response into jq and get coloured, filtered, transformed output without opening a file or writing a parser.TIL
TIL: Forwarding Localhost With ngrok for Webhook DevOriginalOne ngrok command gives any local server a public HTTPS URL. Webhook testing goes from impossible to trivial in under five seconds.TIL
TIL: Testing WebSocket Connections With wscatOriginalA single wscat command opens an interactive WebSocket session to any endpoint. No browser, no boilerplate, just connect and send messages.TIL