Luke Oliff.

The Infrastructure Race Is Changing How We Use AI

·AI·6 min read·Luke Oliff

The biggest story in AI last week wasn’t a model release. It was a compute deal between Anthropic and SpaceX, and that tells you more about where this industry is headed than any benchmark chart.

The Anthropic-SpaceX deal that changed the compute conversation

On May 6, Anthropic announced it had leased the full capacity of SpaceX’s Colossus 1 data center in Memphis, bringing more than 300 megawatts of power and over 220,000 NVIDIA GPUs online. The immediate effect was visible to every Claude user: rate limits doubled, peak-hour throttles disappeared, and API limits for Opus models went up substantially.

The deal is unusual for a few reasons. Colossus 1 was built by xAI for training Grok, Anthropic’s direct competitor. SpaceX acquired xAI in January 2026, and now the same facility that trained a rival model is running Claude inference. The pragmatism is striking: when GPU supply is the binding constraint, you take capacity wherever you can find it.

Anthropic’s CEO Dario Amodei had been public about compute constraints since late 2025. In March 2026 the company introduced peak-hour throttling that reduced Claude Code sessions during US business hours, which hit developers hard. Around 7% of subscribers ran into walls they hadn’t hit before. This deal directly solves that problem. The hardware is online within weeks, and the limits are already relaxed.

The terms are not public, but the scale is. SpaceX’s S-1 filing later in May would reveal that Anthropic agreed to pay $1.25 billion per month for Colossus access through 2029, a total of $45 billion that landed far above earlier analyst estimates.

What rate limit relief means for developers

The immediate impact on anyone using Claude Code was tangible. Developers got doubled five-hour session limits, no peak-hour planning, and higher Opus API rate limits across the board. For anyone building agentic workflows or long-running coding sessions, this was the single biggest improvement in Claude’s usability since the product launched.

But the bigger point is structural. When the AI industry’s compute bottleneck eases, every developer building on these models benefits through cheaper inference, higher rate limits, and fewer 429 errors. The era of rationed access to frontier models where every API call felt precious is ending. The Anthropic-SpaceX deal is a signal that the compute supply curve is about to steepen, and that changes what you can build.

OpenAI ships ads in ChatGPT and new voice API models

The same week, OpenAI started showing ads in ChatGPT. The rollout was quiet, buried in a larger announcement about three new voice API models. The ads appear in free-tier conversations, not Pro or Team accounts, but the direction is clear: the biggest AI companies are under revenue pressure and looking beyond subscriptions.

The new voice API models matter more for the builders reading this. OpenAI expanded its voice offering with lower-latency endpoints and better emotion handling, directly competing with the specialist voice AI providers that have dominated this space. For someone working at a speech-to-text company like I was, this was the week the voice API market became visibly contested in a way it hadn’t been before, with more models, more providers, and tighter latency targets all at once.

The funding numbers tell a different story

Two reports dropped that week that put the scale in perspective. AI venture funding hit $212 billion in 2025, up 85% from $114 billion in 2024. Enterprise spending on generative AI reached $37 billion, a 3.2x increase from the year before. The application layer alone captured $19 billion of that, over 6% of the entire software market.

These numbers are so big they stop being useful for day-to-day decisions. But one detail matters: enterprise application spending grew faster than infrastructure spending. Companies are past the pilot phase and buying AI products that solve specific problems, not just experimenting with models.

What this week taught me about where AI is going

Three things stood out from watching this week as a developer in the space.

First, the rate limit wars are over for now. Compute supply is catching up with demand, and the companies that own their infrastructure will have a structural advantage. The labs that don’t own chips or data centers will pay the margin to those that do.

Second, voice AI is becoming a first-class API surface at every major provider, not just the specialists. OpenAI’s voice model release that week confirmed what had been building for months. TTS and STT are becoming commodity API layers, and the differentiation is moving to latency, emotion control, and pipeline integration.

Third, the gap between frontier labs and everyone else is widening, but not because of model quality. It’s because of compute. Anthropic’s Colossus deal, OpenAI’s chip investments, and Google’s TPU advantage are infrastructure decisions that will determine who can ship what over the next two years. Model benchmarks matter less than who can run them at scale for the lowest cost.

Frequently asked questions

What was the Anthropic-SpaceX Colossus deal?

Anthropic leased the full compute capacity of SpaceX’s Colossus 1 data center in Memphis, covering 300+ megawatts and 220,000 NVIDIA GPUs. The deal immediately doubled Claude Code rate limits and removed peak-hour throttles for Pro and Max subscribers. Orbital AI compute was also mentioned as a long-term interest but has no timeline.

How did OpenAI’s new voice API models affect the market?

OpenAI released three new voice API models the same week it introduced ads in ChatGPT. The models offered lower latency and improved emotion handling, directly competing with specialist voice AI providers.

Why does the rate limit doubling matter for developers?

Higher rate limits mean longer coding sessions, fewer interruptions, and the ability to run more complex agentic workflows without planning around peak hours. It signals that compute capacity is expanding faster than demand, which should lead to lower costs across the industry.

How much did AI venture funding reach in 2025?

AI venture funding reached $212 billion in 2025, up 85% from $114 billion in 2024. Enterprise spending on generative AI hit $37 billion, a 3.2x year-over-year increase.

Is the compute bottleneck really easing?

Yes, but unevenly. The Anthropic-SpaceX deal, OpenAI’s custom chip work, and Google’s TPU infrastructure all point to expanding compute supply. The bottleneck is easing for the major labs that can afford these deals. Independent developers will benefit through lower API prices and higher rate limits.