Two press releases in twenty-four hours
On April 28, 2026, the Wall Street Journal publishes an embarrassing investigation: OpenAI is said to have missed its internal target of one billion weekly users at the end of 2025 (the reported figure topped out at 800 million), missed several monthly revenue targets, and its chief financial officer, Sarah Friar, is said to have privately warned that the company might not be able to fund its compute commitments, whose total runs into the hundreds of billions of dollars. The official response comes back cutting: “pure clickbait”, the company “running at full throttle”, and Altman and Friar declare themselves “totally aligned on buying as much compute as possible.”
The next day, April 29, OpenAI publishes a post with a programmatic title, “Building the compute infrastructure for the Intelligence Age”, and announces a feat in it: the target of 10 gigawatts of secured capacity in the United States, set for 2029 at the launch of Stargate, is already exceeded, “a little over a year later”, including more than 3 gigawatts added in the last 90 days alone.
A damning article on Monday, a triumphant release on Tuesday. The most interesting thing is that both are true. They simply are not talking about the same gigawatt, and telling the two apart is the whole point. Behind OpenAI’s comeback in AI coding, there is an infrastructure bet of unprecedented scale, and to read it correctly you have to learn to distinguish three states of matter for compute: the secured, the built, the plugged-in. The announcement gigawatt is not the one serving your tokens.
One gigawatt, three states
| State | Capacity | Nature of the figure |
|---|---|---|
| Secured (contracts) | > 10 GW | self-reported, April 29 release |
| Under construction | ~9 GW planned, 7 sites | Epoch AI estimate |
| Plugged in (total, end 2025) | 1.9 GW | CFO Sarah Friar statement |
| Plugged in (Abilene, spring 2026) | ~0.3 GW (4 buildings of 8) | Epoch AI estimate + Oracle communication |
Let us start with the secured, since that is what makes the headlines. The 10 gigawatts “exceeded years ahead” are contractual commitments: agreements signed with suppliers to build and deliver capacity. The April 29 release quietly redefined the scope too: “Stargate” no longer designates the joint venture announced in January 2025 with Oracle, SoftBank and MGX, but “OpenAI’s long-term effort to build its compute foundation”, encompassing all of its partners, cloud providers, neoclouds, chipmakers and energy companies. The tasty detail: of the 3 gigawatts secured in 90 days, about 2 come from Amazon Web Services. The joint venture that was to free OpenAI from the hyperscalers now secures its capacity… with them.
This shift is no accident. The Financial Times reported that OpenAI has, in practice, abandoned the joint-venture structure in favor of large bilateral deals, after a year in which the joint venture neither hired nor developed a datacenter of its own, the partners having quarreled over control. The flagship program has had very concrete setbacks: the expansion of the Abilene site beyond 1.2 gigawatt was abandoned in March after financing negotiations failed, Bloomberg citing OpenAI’s “shifting” demand forecasts; Microsoft picked up the adjacent 900 megawatts; and in New Mexico, the gas pipeline meant to power the future giant Project Jupiter site was blocked by the state land office.
Then comes the built: about 9 gigawatts planned across seven US sites, six under construction. Then the plugged-in, and this is where the orders of magnitude contract brutally. According to Epoch AI’s satellite-imagery analysis, cross-checked with an Oracle communication, the flagship Abilene site was operating this spring at around 0.3 gigawatt, four buildings of eight, with outages of several days during its first winter of operation. And OpenAI’s CFO herself put the company’s effective capacity at the end of 2025 at 1.9 gigawatts, all sites and partners combined. Ten gigawatts secured, nine planned, less than two plugged in: there is the honest snapshot of mid-2026.
That leaves the price of the matter. The CEO of Crusoe, the Abilene developer, dropped a figure in front of Stanford students: his datacenters cost $19.2 billion per gigawatt (about €17.7 billion) once you include the gas plants that power them, more than double the $9 to $11 billion of a classic grid-connected “powered shell”. At that rate, the 10 gigawatts secured represent a twelve-figure wall of investment, which restores all their weight to the concerns reported by the WSJ. In our GLM-5.2 report, we cited the Stargate commitment ($500 billion, 10 gigawatts) in a one-to-three-thousand asymmetry against the research compute of a Chinese lab under embargo. This report adds the missing nuance: that half-trillion is a commitment, not a spend, and even less a capacity in service.
Spud, the model sleeping at Abilene
Why this race, when the deployed lags so far behind? Because today’s secured compute is next year’s model, and OpenAI has just demonstrated it with the model that launched its comeback.
GPT-5.5, released on April 23, carried the internal code name “Spud”. Behind the potato hides an engineering event: the first full retraining of OpenAI’s base since GPT-4.5, whose pretraining finished on March 24 at the Abilene Stargate site, on Oracle’s cloud infrastructure and NVIDIA GB200 systems. It is written in black and white in the April 29 release, and the executives did not hide the stakes: Greg Brockman describes “maybe two years of research reaching maturity in this model”, Altman talks of “a new base, a new pretraining”. The trade press even reported that OpenAI shut down its Sora video generator in late March to reallocate compute, a rationing signal the company never officially confirmed.
Now follow the economics of this base, because it explains GPT-5.6’s price range better than any sales pitch. A full pretraining is a capital expenditure: months of cluster, a huge fixed cost, paid once. Everything that follows (the post-trainings, the variants, the tiers) exploits this asset at decreasing marginal cost. GPT-5.6 arrived less than three months after 5.5 with Sol at the same price as 5.5, Terra “competitive with 5.5” at half price and Luna at a sixth: it is the signature of an asset being amortized by declining it, not of a new base being refinanced. An analyst rumor, from a single uncorroborated tweet, goes further by claiming that 5.5 and 5.6 literally share the same “Spud” base of about 4 trillion parameters and that 5.6 closes the series; we report it for what it is, a hypothesis unverifiable to date. But the economic pattern itself is documented by the prices: a generation is paid for in training compute, then monetized in inference tiers.
Selling speed: Cerebras, Spark, Jalapeño
The other half of the compute bet concerns not training but serving, and it lights up an entire part of the economic trade-off described in the first report: speed has become a product.
On January 14, OpenAI signs with Cerebras, the maker of wafer-scale processors, for 750 megawatts of low-latency inference compute delivered by 2028. Announced above $10 billion, the contract was revalued above $20 billion (€18.4 billion) in a Cerebras securities filing in June, with an option to extend to 2 gigawatts by 2030. What this capacity serves is called GPT-5.3-Codex-Spark: a model distilled for speed, served at more than 1,000 tokens per second on the Wafer Scale Engine 3, whose architecture keeps the weights in ultra-fast on-chip memory instead of shuttling them from external memory. And in late June, the GPT-5.6 preview promised Sol “up to 750 tokens per second” on Cerebras during July, a prospective figure, announced by OpenAI alone, which we will re-verify.
Then there is the in-house chip. On June 24, OpenAI unveiled Jalapeño, its first inference ASIC Application-Specific Integrated Circuit. A chip designed for a single task, wired in hard silicon. Very power-efficient at what it was built for, incapable of the rest. , co-designed with Broadcom, fabricated at TSMC, from spec to silicon in nine months. Engineering samples already run Spark in the lab; commercial deployment is targeted for late 2026, with 10 gigawatts of in-house accelerators on the 2029 horizon.
| Partner | Scope | Status mid-2026 |
|---|---|---|
| Cerebras | 750 MW low-latency, > $20B (€18.4B) | signed (SEC filing), serves Spark |
| Broadcom / TSMC | in-house Jalapeño chip, 10 GW targeted by 2029 | lab samples, deployment late 2026 |
| NVIDIA | 10 GW letter of intent, up to $100B (€92B) | "will not proceed as announced" |
| AMD | 6 GW MI450 + warrant for ~10% of equity | first phase late 2026 |
| AWS | $38B (€35B) over 7 years | signed, ~2 GW secured in spring |
| Microsoft Azure | additional $250B (€230B) commitment | primary cloud partner until 2032 |
Who pays whose subscription
Because this is the paradox of the comeback: the conquest of developers is happening at prices the costs do not yet justify.
SemiAnalysis ran the most telling experiment of the spring: buy each subscription tier from both providers and exhaust them methodically with agentic coding workloads. Result: a ChatGPT Pro subscription at $200 a month (€184) fully used delivers about $14,000 of tokens valued at the API rate (€12,900); the equivalent at Claude, Max 20x, about $8,000 (€7,400). A lever of 40 to 70 times between the price paid and the value consumed. The figure needs its methodological nuance: value at the API rate includes the provider’s margin, and the real compute cost is more like 15 to 20 times the subscription. But even corrected, the verdict holds: intensive users are massively subsidized. SemiAnalysis estimates that OpenAI goes into negative gross margin as soon as a Plus subscriber uses more than 11.4% of their cap, and to zero around 5.7% on the high end; Anthropic, better protected by its dedicated chips, would hold out to about 20 and 10% respectively.
To understand why it happened now, you have to look at what an agent task actually consumes. SemiAnalysis puts typical agentic work at about 96,000 tokens, up to a thousand times a classic chat prompt: where a conversation question costs a fraction of a cent, a coding session that reads the repository, reasons, tests and corrects itself mobilizes the equivalent of a small book at each loop iteration. Subscriptions were priced for the chat user, whose serving cost runs in tens of cents a month (about $0.70 for a free user, per the same analysis), and Sacra estimates that a third of OpenAI’s inference costs still goes to soaking up that free tier. Then agents arrived in the same plans, with consumption profiles a thousand times higher. Telecom operators experienced exactly this collision: “unlimited” plans sized for voice and mail, hit by video streaming, three orders of magnitude of traffic under the same contract. All of the subscription spring told in the first report (the split announced then canceled at Anthropic, the Uber case, the quotas going up and down) flows from this collision between a pricing designed for conversation and a usage gone industrial.
The overall finances tell the same story at scale. According to Sacra and The Information data, OpenAI posted about 33% gross margin in 2025, with inference costs of $8.4 billion (€7.7 billion) projected to $13 billion in 2026, cash burn on the order of $27 billion (€25 billion) this year, and no positive cash flow expected before 2030. On the other side: 50 million paying subscribers and 900 million weekly users claimed in February, annualized revenue of about $25 billion (€23 billion) mid-2026, and an IPO filing submitted confidentially on June 8, for a valuation floated as high as $1 trillion. Sam Altman had set the tone as early as January 2025, with a now-famous candor: “we are currently losing money on Pro subscriptions”, because people “use it much more than we expected”. The bet that makes all of this tenable carries a figure, put forward by Sarah Friar: between GPT-5 and GPT-5.4, the serving cost of a given capability level is said to have fallen by about 97%. As long as this deflation curve holds, subsidizing today’s usage amounts to buying the customer base that will pay tomorrow’s collapsed costs. If it stalls, the S-1 will tell a different story.
The other model: Anthropic and multi-cloud
The contrast with Anthropic deserves its section, because the two compute strategies are almost ideal types. Where OpenAI contracts gigawatts in every direction under a single brand, Anthropic embraces an arbitrage multi-cloud: “we train and serve Claude on AWS Trainium, Google TPU Tensor Processing Unit. Google's ML acceleration ASIC. Not a lone chip but a network: chips linked directly by the ICI interconnect in a torus topology, each combining TensorCores (dense matmul) and SparseCores (embeddings, collectives). It only knows how to talk to the XLA compiler. and NVIDIA GPUs”, each workload routed to the most suitable chip. The figures are of the same order of excess, but the split differs: the Project Rainier campus in Indiana, about half a million Trainium2 chips in service since the fall, an AWS deal extended to 5 gigawatts and more than $100 billion over ten years; up to a million Google TPUs and more than a gigawatt online in 2026, extended in April with Broadcom toward 3.5 additional gigawatts from 2027; a first gigawatt of NVIDIA systems; a $30 billion Azure commitment. And one telling one-off deal: more than 300 megawatts taken over from SpaceX in May, which Anthropic explicitly made the fuel for doubling Claude Code’s limits, one of the episodes of the subscription spring. The capacity bought reads directly in the quotas.
Beware, however, of the scope trap, because it lurks on both sides. Cleanview’s satellite analysis pointed out that Anthropic’s Indiana campus passed a gigawatt in service in March, “five times Stargate’s capacity at the same moment”: the formula is accurate, but it compares a campus to a program, and different states of matter for compute. The honest symmetry fits in one sentence: Anthropic plugged in gigawatts faster, OpenAI contracted more of them. Which of the two is right depends on a single variable, the speed at which inference demand will catch up with the construction sites, and no one knows it. What already measures, on the other hand: annualized revenue up from $9 billion at the end of 2025 to more than $30 in April then $47 in May, the sector’s strongest acceleration, 80% driven by enterprises.
GPT-6, or the next base
That leaves the horizon, and it demands sobriety, because folklore has already done damage there. What is established: GPT-6 will arrive faster than the 28 months that separated GPT-4 from GPT-5, and Altman has named its priority axis, memory (“people want memory”). What has already been contradicted by the facts: the “40%” leap on benchmarks, the two-million-token context, the April launch, the super-app with a built-in browser (irony: the merger did happen, in July, and the Atlas browser is closing). What comes from a single tweet: GPT-5.6 last of the 5.x series, GPT-6 “within a month”, on a base much bigger than Spud. And what comes from the market: the Polymarket bets on a 2026 release, a sentiment, not information. We froze this report on July 11; the update box is planned, as on the two previous installments of this series.
At bottom, though, one thing is almost certain, and it sums up this whole report: the next base will be paid for in gigawatts, and it will train on sites that, today, are construction yards. It is the sector’s grammar that got written this spring: models pass, benchmarks go obsolete, prices readjust every week, but the gigawatts stay, and the frontier belongs to whoever can fund the next pretraining. To that grammar, our GLM-5.2 report had documented the exception that contests it: a lab under embargo manufacturing the frontier on the cheap, unable to buy it. Between the half-trillion contracted on one side and the $216 million of research compute on the other, the 2027 race will play out less on the intelligence of the models than on the return on the capital that brings them into being. And on that ground, no one has yet proven that spending more means lasting longer.
Sources and method
Currencies converted at the indicative rate of 1 USD = 0.92 EUR (mid-2026). Editorial freeze date: July 11, 2026; an update box will be added if the facts change (notably: Sol served on Cerebras, status of Anthropic pricing). Labels: verified fact (primary source), estimate (third-party analysis), self-reported (vendor figure), hypothesis (rumor or assumed reasoning).
Stargate: secured, built, plugged in
- Self-reported: OpenAI, Building the compute infrastructure for the Intelligence Age (openai.com, April 29, 2026): > 10 GW secured, +3 GW in 90 days, “ecosystem” scope; GPT-5.5 “trained on our flagship Abilene site” (OCI, NVIDIA GB200). Bloomberg (April 29-30): ~2 of the 3 GW come from AWS.
- Fact: Wall Street Journal (April 28, 2026): missed user and revenue targets, funding concerns; OpenAI’s public responses the same day. Financial Times: abandonment of the joint-venture structure in favor of bilateral deals.
- Estimate: Epoch AI, OpenAI Stargate: where the US sites stand (satellite imagery, spring 2026): ~0.3 GW in service at Abilene (4/8 buildings), ~9 GW planned. Fact: effective capacity at the end of 2025 of 1.9 GW (Sarah Friar statement, via Data Center Dynamics).
- Fact: Bloomberg (March 6, 2026): Abilene expansion abandoned, 900 MW picked up by Microsoft; Project Jupiter pipeline blocked (New Mexico, March). Cost per GW: statement by Crusoe’s CEO ($19.2B/GW gas included, against $9-11B in powered shell), via Cleanview/Distilled.
Spud and the economics of bases
- Fact: “Spud” = GPT-5.5’s code name (Axios, April 23, 2026), first full retraining of the base since GPT-4.5, pretraining finished March 24 (The Information); Brockman (“two years of research”) and Altman (“a new base”) verbatim. GPT-5.6 prices: openai.com/index/gpt-5-6.
- Estimate: Sora shutdown in late March for compute reallocation (trade press, not confirmed by OpenAI).
- Hypothesis: shared 5.5/5.6 base “~4T”, “5.6 last of the series”: single analyst tweet (~July 7), uncorroborated, reported as such.
Speed as a product
- Fact: Cerebras partnership (January 14, 2026, 750 MW through 2028) revalued > $20B by SEC filing (Reuters, June 23); GPT-5.3-Codex-Spark on WSE-3 (> 1,000 tok/s, Pro preview, adjustable limits). Sol “up to 750 tok/s” on Cerebras: OpenAI preview page (June 26), prospective, with no dedicated Cerebras release.
- Fact: Jalapeño: OpenAI/Broadcom announcement (June 24, 2026, via Axios/CNBC); Broadcom memorandum of understanding 10 GW (October 2025). NVIDIA: 10 GW / $100B letter of intent (September 2025), “will not proceed as announced” (The Information). AMD: 6 GW + warrant (October 2025). AWS: $38B/7 years. Azure: $250B additional commitment.
Subscription economics
- Estimate: SemiAnalysis (June 2026, subscription exhaustion tests): ~$14,000 of token value on ChatGPT Pro $200, ~$8,000 on Claude Max 20x, margin thresholds (11.4% / 5.7%; ~20% / ~10%), agentic task ~96,000 tokens, free user ~$0.70/month. API-value ≠ compute-cost nuance (~15-20×) spelled out in the body of the text.
- Estimate: Sacra: gross margin ~33% (2025), inference $8.4 → 14.1B, cash burn ~$27B (2026), positive cash flow not before 2030; ~1/3 of inference costs allocated to the free tier.
- Fact: 900M weekly users and 50M paying subscribers (February 27 announcement, via TechCrunch); confidential S-1 (June 8); “we are currently losing money on openai pro subscriptions” (Altman, January 2025); ~97% serving-cost deflation GPT-5 → 5.4 (Sarah Friar, All-In podcast).
Anthropic
- Fact: claimed multi-cloud (Trainium, TPU, NVIDIA); Project Rainier (~500,000 Trainium2, extension to 5 GW, > $100B/10 years); Google TPU (up to 1M chips, > 1 GW in 2026) + Broadcom extension (~3.5 GW, 2027); 1 GW NVIDIA (8-K); $30B Azure; SpaceX deal > 300 MW tied to raising Claude Code limits (May 6 release). Annualized revenue $9 → 30 → 47B (Anthropic statements, Feb-May 2026).
- Estimate: Indiana campus > 1 GW in March, “5× Stargate at the same moment” (Cleanview): asymmetric scopes, flagged in the body of the text.
GPT-6
- Fact: tightened schedule and memory priority (Altman public statements). Hypothesis: everything else (specs, dates, base), including the already-contradicted folklore, listed in the body of the text so it is not picked up.
Method note. This report systematically distinguishes three states of compute capacity (secured, built, plugged in) because their confusion is the main vector of distortion for the sector’s figures, in both directions: it manufactures both illusory triumphs and illusory fiascos. The “secured” amounts are contractual commitments, not realized spending; the private financial figures (margins, cash burn) come from third-party analysts and not audited accounts, OpenAI’s S-1 remaining confidential as of the freeze date. This article is the English adaptation of a French original, published on July 16, 2026.