Back to Home

May 2026

From: Brian and Gaby

Subject: Compute strategy is the product strategy

Compute Strategy Is Product Strategy

The shift from subsidy to scarcity in AI is fundamentally a compute story.

Driving this change are two interconnected dynamics. One, token demand accelerated faster than expected, driven by the agentic explosion that kicked off in late 2025. Second, the labs need to grow up financially. Their revenue is scaling at a blistering, never-before-seen pace, but so are their costs. And with IPOs looming later this year, they need to present something more durable than growth at any price.

In the subsidy phase, usage was the scoreboard. More users, more tokens, more workloads, more proof that the market was real. But public-market investors (and the public in general) will ask a harder question: where’s the profit? Once that becomes the question, compute strategy, takes on supreme significance.

By compute strategy, we mean how each company sources compute, prices access to it, allocates it across products and customers, improves utilization, and turns scarce infrastructure into margin advantage. These companies monetize tokens, which are software wrapped around compute. They also depend on compute to build their products, i.e. train their models, and innovate internally. The allocation of that resource is becoming more than a balance sheet exercise, but a management skill that defines the ceiling of these companies.

Every company is playing this field differently. As we watch the unfolding competition in AI, we increasingly apply the lens that compute strategy is upstream of product strategy, and often upstream of company strategy itself. The next phase of AI competition will be won by who can turn compute into a compounding system.

In this newsletter, we want to assess some of the leading AI companies through the lens of their compute strategy because we believe it reveals more about their prospects than their current revenue run rate or latest model benchmark.

Let’s dig in.

An overview of the field we’re covering

SpaceX

Six months ago, SpaceX was not a compute company. It was a rocket company, a satellite company, and an internet company. But perhaps Elon’s real strategy may have been clear all along – 

Last week, as SpaceX went public, the story was all about AI infrastructure: large-scale GPU capacity with terrestrial data centers and orbital compute becoming increasingly concrete, and the capital required to execute on this vision. 

SpaceX knows how to build hard physical systems quickly. It knows how to site, power, cool, deploy, and operate complex infrastructure. SpaceX, as driven by Elon’s pace, knows how to compress timelines in markets where acceleration was once imagined to be impossible. 

Colossus made this obvious. xAI did not just buy GPUs; it built a giant compute facility at remarkable speed. He built Colossus 1, which included 100K H100s using 300 MW of power, in 122 days, doubling its capacity 92 days later. Jensen himself said, “What Elon and the team achieved is singular. It has never been done before,” given his expertise in engineering, construction, and marshalling resources towards a goal.

Now, Elon is on a mission to manufacture 1 terawatt of processors annually at Terafab, and launch 100 GW of compute into space annually with SpaceX in the near future. The dream? A petawatt in space. That is ~800x more power than the entire US power capacity today. 

And when that compute was underutilized given Grok’s underperformance in market, Elon quickly filled Colossus with willing and eager tenants in Cursor, Google, and Anthropic. Google is paying nearly $1B a month as a tenant, with Anthropic paying $1.25B on top. And just as we’re writing this – SpaceX landed another tenant in Reflection AI, paying $150M per month for access to Colossus 2. 

SpaceX’s strategy is therefore not simply to “own a lot of GPUs.” It is to build the compute factories, secure the power, create or acquire the product surfaces that can consume that compute, and tell investors that the company is not only launching rockets but operating the industrial base for AI.

OpenAI

OpenAI has its mojo back. 

In May, the vibe shifted. You could hear it in conversations with engineers, founders, and platform teams who were transitioning from Claude devotees Codex fans.

Under the hood, a compute advantage was driving a product win. Codex was not only good; it was available and persistent whereas Claude had excessive downtime, unexpected outages, and lots of spinning wheels. OpenAI took off through the agentic usage curve because they had the compute capacity. Their strategy for compute is represented in their product, especially as it compares to Claude Code. Codex moves more work into cloud-delegated, parallel workflows where OpenAI can manage, meter, and prioritize scarce inference directly. Codex feels more like a handoff, which gives OpenAI room to decide how much cloud inference to spend, when to run tasks remotely, how to prioritize agentic usage, and how to package that into paid ChatGPT plans. 

OpenAI’s broader compute strategy remains the most aggressive in the market. Stargate is the public expression of it, inking deals across HBM, DRAM, networking, power, and data center capacity to make those accelerators productive.

For some quick math: if a power user consumes ~1M tokens per day, and serving that user for a year takes roughly 60W of sustained H100-generation compute at realistic utilization, then each H100 can support around a dozen such users. OpenAI's ~1.7M H100 equivalents versus Anthropic's ~1M leaves a 700,000-GPU advantage — enough headroom to serve on the order of 8 million additional power users. Put differently, OpenAI's compute lead alone could host an entire Anthropic-scale user base several times over before either company touches a new chip.

The value of compute compounds. OpenAI has been aggressive in how they allocate compute, pushing it when they want to make one move (like consumer app Sora) and pulling back when that move doesn’t prove to be the right one (shutting Sora down). It’s a constant push and pull, but one that OpenAI is better positioned for given the volume of compute they have on hand and how much they’re bringing online in the next 1-3 years. 

Anthropic

Anthropic sits on the other side of this trade. The company has generally positioned itself as the more disciplined frontier lab: enterprise-first, safety-forward, less theatrical, and more careful about the financial risks of the compute race. xAI and SpaceX are almost the opposite cultural animal: speed, scale, spectacle, and infrastructure audacity. And yet in a scarce compute market, ideology bends toward capacity.

Anthropic is in pole position in many ways in the race right now – more effective at using compute (training and deploying equally performant models with less compute), capturing a higher margin (reports of 40-70%, assumed to be higher than OpenAI), and valued higher in the market ($965B vs OpenAI’s $852B). But those in pole position don’t always win the race. 

Anthropic’s stance on compute has been less aggressive than OpenAI’s. During an episode of the Dwarkesh Podcast in February 2026, Dario spoke quite flippantly about his perspective on overbuying compute – “I could assume that revenue will continue growing 10x a year. So it will be $100B in 2026 and $1T in 2027. And I’d buy $1T of compute that starts at the end of 2027. And if my revenue is not $1T, even $800B, there is no force on earth, no hedge on earth, that would stop me from going bankrupt if I buy that much compute”. This directly impacted Anthropic’s product. Check out the uptime of Claude Code from December through February, followed by that of the last three months. Nearly every week in March experienced a major outage. The compute constraint, combined with a speed of product deployment that frankly couldn’t handle the usage numbers they were seeing, resulted in an experience we all probably remember. 

Many things make the Colossus relationship....um, interesting. First, Musk love shitposting on Anthropic and one could not imagine stranger bedfellows than him and Dario. But Anthropic has an incredible advantage with their ability to use compute flexibility and get superior gross margins to other labs.

Anthropic’s compute posture has been about procuring capacity within reason, and ensuring the flexibility of their compute. Anthropic has been explicit that Claude trains and runs flexibly across AWS Trainium, Google TPUs, and NVIDIA GPUs, with incredible flexibility across where any single workload can run.

Portability across chips gives Anthropic more control over its margins, its supply chain, and its product roadmap. It lets the company match workloads to hardware instead of forcing every job through the same constrained accelerator market. It gives Anthropic leverage with suppliers, reduces dependence on any single ecosystem, and allows the company to optimize differently for training, batch inference, latency-sensitive serving, and enterprise deployments. In a world where NVIDIA remains dominant but supply, cost, and power are all constraints, being portable across Trainium, TPUs, and GPUs is their business advantage. 

This is the point of the piece in miniature: compute strategy drives business strategy. Anthropic’s appreciation for the true cost of compute appears to have pushed it toward the right business model earlier. Enterprise usage is less fickle than consumer usage, has better contractual visibility, supports higher willingness to pay, and gives the company a more rational basis for capacity commitments. The constraint forced focus, and that focus turned into a differentiated market position.

Google

Google is in its own league. It is the one vertical integrator in AI. It has the models, the chips, the cloud, the products, the operating system, the browser, the enterprise suite, the research lab, and the distribution. 

Part of this is just the mere fact that they have been in the game far longer. In 2022, Google was posting about investing $9.5B into data centers across the US, a comically small number in comparison to today’s investments. They had a rapidly scaling ads and search business that required a massive footprint (for its time) of data centers. Google had land, power, permits, and more that put them far ahead of their competitors entering the AI race in the 2022/2023 time frame. Of course, net new investment was needed, but they were playing with a different set of physical assets, internal knowledge on operating data centers at scale, and an internal chip team with several tapeouts under their belts.

The TPU is the center of their strategy in today’s AI era. Ironwood, Google’s seventh-generation TPU, is built for large-scale training and inference. This is the hardware behind Gemini, Nano Banana, and Google’s broader AI stack. It is also a TPU that is being sold widely beyond Google. Given the TPU’s prominence outside of Google, the company has been in a delicate balance – what compute should be allocated to Google, to DeepMind, to Anthropic, to Meta, to inference, and to training? The TPU has historically lowered costs for Google’s largest workloads – search and ads. But, with a look towards AI, the TPU brings a new level of complexity to Google’s infrastructure stack. The TPU matters most because Google does not need to rent its entire destiny from NVIDIA. It can use NVIDIA where useful, the TPU where most effective, then on-device models for workloads suitable for their hardware

Google’s on-device AI gives the company a path to run smaller, more private, lower-latency models close to the user, while larger Gemini models and cloud inference handle heavier tasks. In a world where every task should not pay the cost of a frontier cloud model, deciding what runs on-device, what runs in the cloud, and what moves between them becomes a core product architecture decision.

Google’s challenge is not whether it has the pieces. It has more pieces than anyone. The question is whether vertical integration becomes product clarity. Google has to protect Search, grow Cloud, defend Android, monetize Workspace, serve developers, compete with OpenAI and Anthropic, and avoid breaking the economics of its core business. The infrastructure is world-class; the product story still has to become as coherent as the compute story.

What this means for startups

It is easy to look at SpaceX, OpenAI, Anthropic, and Google and conclude that compute is only a game for giants. We think the opposite: when the largest companies in the world reorganize themselves around a new infrastructure layer, startups get created in the gaps.

In past technology cycles, those gaps themselves were not large enough to sustain companies. In the inference era, with the scale ceiling moving higher every day, these gaps are multi-billion, if not multi-trillion, dollar opportunities.  

This is the core of our compute thesis. The demand for compute is forcing a massive capex buildout across silicon, power, memory, networking, cooling, data centers, and software. Fundamental bottlenecks emerge, and incumbents get pushed beyond incremental improvement. From this, new opportunities open for founders taking first-principles approaches to how compute is produced and scaled.

The giants will continue on their course, gobbling up compute as they go. Alongside them, a new world of startups will enter the race, and we’re excited to back more of them. 

As always, if you know someone building in this category, or an expert we should meet, let us know.

P.S. In May, we launched The Compute 100 and a daily newsletter, highlighting the 50 leading public and 50 leading private companies in compute. We’ve since expanded it into a podcast featuring conversations with the founders, operators, and investors building the next era of compute. Give it a listen and let us know what you think – and if there’s a company we’re missing, we’d love to hear about it.

Best,
Brian & Gaby