June 1, 2026 12 minutes min read

AI Infrastructure Index

Tracking the key metrics defining the AI infrastructure landscape — compute capacity, GPU deployments, and hyperscaler capex. Updated weekly.

AI Infrastructure Index

A weekly temperature check of the global AI infrastructure landscape. We track compute capacity, cloud capital expenditure, data center construction, and energy infrastructure — delivering structured weekly observations on the AI compute ecosystem.

Cover image: Self-provided

Weekly Overview

June 1-7, 2026: AI infrastructure investment and deployment entered an accelerated phase. NVIDIA Blackwell B200 GPU shipments hit a record high, Big Three cloud capex continued to climb, and inference compute demand surpassed training for the first time. Key metrics at a glance:

Metric This Week MoM Change YTD
NVIDIA B200 Monthly Shipments 1.2M units +33% 4.1M units
Big 3 Cloud AI Capex (Q2 est.) $58B +85% YoY $112B
Inference Share of Total Compute 52% +8pp
New Data Centers Under Construction 47 +5 312
AI Data Center Power Demand (est.) 62 GW +9 GW
AI Chip Startup Funding $1.23B +$410M $4.78B

1. Compute Supply: Blackwell Ramp

1.1 GPU Shipments

NVIDIA's Blackwell B200 reached 1.2 million units shipped in May, up 33% from April's 900,000. TSMC's CoWoS-L advanced packaging capacity expanded to 450,000 12-inch equivalent wafers per month, remaining the most constrained link in the supply chain:

Package Type Monthly Capacity (12" eq.) Utilization Bottleneck
CoWoS-L (Blackwell) 450K wafers 98% Severe
CoWoS-S (H100/B200) 320K wafers 95% Moderate
InFO (Edge AI) 280K wafers 82% Light
SoIC (3D Stacking) 120K wafers 91% Moderate

The CoWoS-L capacity gap stands at approximately 15-20% of demand. NVIDIA has placed additional reservations with TSMC for H2 2026, targeting 650K wafers per month by year-end.

1.2 Competitive Landscape

AMD MI400 shipped approximately 180,000 units in May, primarily serving Microsoft and Meta internal clusters. Google's TPU v7 entered large-scale deployment in May, with a planned 1 million+ unit deployment by end of 2026. AWS Trainium3 silicon has been delayed to Q3 2026 due to design change respins.

Chip May Shipments Q2 Cumulative Key Customers
NVIDIA B200 1.2M 2.7M All hyperscalers + enterprise
AMD MI400 180K 380K Microsoft, Meta
Google TPU v7 Internal deployment 300K Google internal
AWS Trainium3 Delayed to Q3 AWS internal

2. Cloud Capex: Hyperscaler Arms Race

2.1 Q2 2026 Estimated CapEx

Combined AI capital expenditure from the Big Three cloud providers is estimated at $58 billion for Q2 2026, up 85% YoY. Over 70% goes to servers and GPU procurement, with the remainder for data center infrastructure:

Provider Q2 AI CapEx Est. YoY Growth Primary Investment
Microsoft Azure $23B +92% Global data center expansion, OpenAI clusters
Amazon AWS $19B +78% Trainium infrastructure, Bedrock inference
Google Cloud $16B +84% TPU v7 deployment, Gemini cluster expansion

2.2 Data Center Construction

Large-scale AI data centers (>100MW) under construction or planned reached 47 in May, adding 5 new facilities. Geographic distribution: North America (28), Europe (9), Asia-Pacific (10). A notable trend is the rapid adoption of liquid cooling — 78% of new facilities use direct liquid cooling (DLC) or immersion cooling, up from 52% in 2025.

3. Inference vs. Training: Structural Shift

3.1 Inference Overtakes Training

API usage data from OpenAI and Anthropic shows inference compute surpassed training compute in May for the first time, reaching 52% of total compute demand. This inflection point arrived 6-9 months earlier than industry expectations, driven by:

  • Combined ChatGPT/Claude MAU exceeding 800 million
  • Explosive growth in AI agent automation inference calls
  • Accelerated enterprise inference deployment (RAG, document analysis, customer support)
Compute Category Share MoM Growth YTD Share
Inference 52% +18% 41%
Training 38% +7% 47%
R&D/Testing 10% +5% 12%

3.2 Inference Hardware Trends

The inference demand explosion is reshaping procurement patterns. NVIDIA estimates inference will account for 60% of its data center GPU shipments by end of 2026. Meanwhile, dedicated inference chip orders (Groq LPU, Cerebras Wafer-Scale, D-Matrix) grew 140% QoQ in Q2.

4. Energy & Power: AI's Growing Appetite

4.1 Power Consumption Overview

Global AI data center power demand reached 62 GW operational in May, with annualized consumption of approximately 543 TWh — about 1.8% of global electricity generation. By 2027, AI data center power demand is projected to reach 150 GW, exceeding Sweden's total national consumption.

Region AI DC Power (GW) Share of Regional Generation Primary Energy Source
United States 28.5 GW 3.2% Natural gas + nuclear + renewables
Europe 14.2 GW 1.9% Nuclear + wind
Asia-Pacific 16.8 GW 1.5% Coal + natural gas
Other 2.5 GW 0.3% Mixed

4.2 Energy Solutions

Microsoft's agreement with Constellation Energy to restart Three Mile Island Unit 1 (835 MW) has entered federal regulatory review, targeting 2028 operation. Google's power purchase agreement with Kairos Power for small modular reactors (6 x 50 MW) has initiated site selection. AWS continues large-scale wind and solar procurement, investing $4.5 billion in renewables in 2026.

5. This Week's Key Events

5.1 OpenAI in Talks with Oracle for Mega-Cluster

OpenAI is reportedly in negotiations with Oracle to build a 2-million GPU training cluster, with estimated total investment exceeding $50 billion. If built, it would be 5x the size of the largest known single training cluster.

5.2 BIS Tightens AI Chip Export Controls

On June 2, the Bureau of Industry and Security published revised export controls, lowering the compute density threshold from 200 PFLOPS to 150 PFLOPS, further restricting AI compute access for non-allied nations. NVIDIA stated the impact on B200 shipments would be limited, as primary customers are in the US and allied nations.

5.3 Anthropic Renews $4B Compute Contract with CoreWeave

Anthropic signed a 5-year, $4 billion compute lease agreement with GPU cloud provider CoreWeave to secure next-generation Claude training and inference capacity. CoreWeave also announced its 12th data center in Norway, leveraging Nordic hydropower to reduce carbon footprint.

6. AI Chip Startup Funding

AI chip startups raised $1.23 billion in May. Key transactions:

Company Round Amount Core Technology
Groq Series E $450M LPU inference chip, $8.5B valuation
MatX Series C $280M Edge AI inference SoC
d-Matrix Series D $250M Digital in-memory compute inference chip
Etched Series B $150M Transformer-specific ASIC
Tenstorrent Series D $100M RISC-V AI accelerator

Next Week's Watchlist

  1. NVIDIA GTC Taipei 2026 (June 10) — potential Rubin architecture details
  2. FERC review progress on AI data center grid connection applications
  3. TSMC June CoWoS-L capacity expansion target
  4. OpenAI / Oracle mega-cluster contract announcement
  5. EU AI Act energy efficiency requirements for large-scale training clusters

Disclaimer: Data and forecasts in this weekly report are based on public information and industry estimates. POC.HK strives for accuracy, but some forward-looking indicators are model-based estimates and may deviate from actual figures.