A weekly temperature check of the global AI infrastructure landscape. We track compute capacity, cloud capital expenditure, data center construction, and energy infrastructure — delivering structured weekly observations on the AI compute ecosystem.
Cover image: Self-provided
Weekly Overview
June 1-7, 2026: AI infrastructure investment and deployment entered an accelerated phase. NVIDIA Blackwell B200 GPU shipments hit a record high, Big Three cloud capex continued to climb, and inference compute demand surpassed training for the first time. Key metrics at a glance:
| Metric | This Week | MoM Change | YTD |
|---|---|---|---|
| NVIDIA B200 Monthly Shipments | 1.2M units | +33% | 4.1M units |
| Big 3 Cloud AI Capex (Q2 est.) | $58B | +85% YoY | $112B |
| Inference Share of Total Compute | 52% | +8pp | — |
| New Data Centers Under Construction | 47 | +5 | 312 |
| AI Data Center Power Demand (est.) | 62 GW | +9 GW | — |
| AI Chip Startup Funding | $1.23B | +$410M | $4.78B |
1. Compute Supply: Blackwell Ramp
1.1 GPU Shipments
NVIDIA's Blackwell B200 reached 1.2 million units shipped in May, up 33% from April's 900,000. TSMC's CoWoS-L advanced packaging capacity expanded to 450,000 12-inch equivalent wafers per month, remaining the most constrained link in the supply chain:
| Package Type | Monthly Capacity (12" eq.) | Utilization | Bottleneck |
|---|---|---|---|
| CoWoS-L (Blackwell) | 450K wafers | 98% | Severe |
| CoWoS-S (H100/B200) | 320K wafers | 95% | Moderate |
| InFO (Edge AI) | 280K wafers | 82% | Light |
| SoIC (3D Stacking) | 120K wafers | 91% | Moderate |
The CoWoS-L capacity gap stands at approximately 15-20% of demand. NVIDIA has placed additional reservations with TSMC for H2 2026, targeting 650K wafers per month by year-end.
1.2 Competitive Landscape
AMD MI400 shipped approximately 180,000 units in May, primarily serving Microsoft and Meta internal clusters. Google's TPU v7 entered large-scale deployment in May, with a planned 1 million+ unit deployment by end of 2026. AWS Trainium3 silicon has been delayed to Q3 2026 due to design change respins.
| Chip | May Shipments | Q2 Cumulative | Key Customers |
|---|---|---|---|
| NVIDIA B200 | 1.2M | 2.7M | All hyperscalers + enterprise |
| AMD MI400 | 180K | 380K | Microsoft, Meta |
| Google TPU v7 | Internal deployment | 300K | Google internal |
| AWS Trainium3 | Delayed to Q3 | — | AWS internal |
2. Cloud Capex: Hyperscaler Arms Race
2.1 Q2 2026 Estimated CapEx
Combined AI capital expenditure from the Big Three cloud providers is estimated at $58 billion for Q2 2026, up 85% YoY. Over 70% goes to servers and GPU procurement, with the remainder for data center infrastructure:
| Provider | Q2 AI CapEx Est. | YoY Growth | Primary Investment |
|---|---|---|---|
| Microsoft Azure | $23B | +92% | Global data center expansion, OpenAI clusters |
| Amazon AWS | $19B | +78% | Trainium infrastructure, Bedrock inference |
| Google Cloud | $16B | +84% | TPU v7 deployment, Gemini cluster expansion |
2.2 Data Center Construction
Large-scale AI data centers (>100MW) under construction or planned reached 47 in May, adding 5 new facilities. Geographic distribution: North America (28), Europe (9), Asia-Pacific (10). A notable trend is the rapid adoption of liquid cooling — 78% of new facilities use direct liquid cooling (DLC) or immersion cooling, up from 52% in 2025.
3. Inference vs. Training: Structural Shift
3.1 Inference Overtakes Training
API usage data from OpenAI and Anthropic shows inference compute surpassed training compute in May for the first time, reaching 52% of total compute demand. This inflection point arrived 6-9 months earlier than industry expectations, driven by:
- Combined ChatGPT/Claude MAU exceeding 800 million
- Explosive growth in AI agent automation inference calls
- Accelerated enterprise inference deployment (RAG, document analysis, customer support)
| Compute Category | Share | MoM Growth | YTD Share |
|---|---|---|---|
| Inference | 52% | +18% | 41% |
| Training | 38% | +7% | 47% |
| R&D/Testing | 10% | +5% | 12% |
3.2 Inference Hardware Trends
The inference demand explosion is reshaping procurement patterns. NVIDIA estimates inference will account for 60% of its data center GPU shipments by end of 2026. Meanwhile, dedicated inference chip orders (Groq LPU, Cerebras Wafer-Scale, D-Matrix) grew 140% QoQ in Q2.
4. Energy & Power: AI's Growing Appetite
4.1 Power Consumption Overview
Global AI data center power demand reached 62 GW operational in May, with annualized consumption of approximately 543 TWh — about 1.8% of global electricity generation. By 2027, AI data center power demand is projected to reach 150 GW, exceeding Sweden's total national consumption.
| Region | AI DC Power (GW) | Share of Regional Generation | Primary Energy Source |
|---|---|---|---|
| United States | 28.5 GW | 3.2% | Natural gas + nuclear + renewables |
| Europe | 14.2 GW | 1.9% | Nuclear + wind |
| Asia-Pacific | 16.8 GW | 1.5% | Coal + natural gas |
| Other | 2.5 GW | 0.3% | Mixed |
4.2 Energy Solutions
Microsoft's agreement with Constellation Energy to restart Three Mile Island Unit 1 (835 MW) has entered federal regulatory review, targeting 2028 operation. Google's power purchase agreement with Kairos Power for small modular reactors (6 x 50 MW) has initiated site selection. AWS continues large-scale wind and solar procurement, investing $4.5 billion in renewables in 2026.
5. This Week's Key Events
5.1 OpenAI in Talks with Oracle for Mega-Cluster
OpenAI is reportedly in negotiations with Oracle to build a 2-million GPU training cluster, with estimated total investment exceeding $50 billion. If built, it would be 5x the size of the largest known single training cluster.
5.2 BIS Tightens AI Chip Export Controls
On June 2, the Bureau of Industry and Security published revised export controls, lowering the compute density threshold from 200 PFLOPS to 150 PFLOPS, further restricting AI compute access for non-allied nations. NVIDIA stated the impact on B200 shipments would be limited, as primary customers are in the US and allied nations.
5.3 Anthropic Renews $4B Compute Contract with CoreWeave
Anthropic signed a 5-year, $4 billion compute lease agreement with GPU cloud provider CoreWeave to secure next-generation Claude training and inference capacity. CoreWeave also announced its 12th data center in Norway, leveraging Nordic hydropower to reduce carbon footprint.
6. AI Chip Startup Funding
AI chip startups raised $1.23 billion in May. Key transactions:
| Company | Round | Amount | Core Technology |
|---|---|---|---|
| Groq | Series E | $450M | LPU inference chip, $8.5B valuation |
| MatX | Series C | $280M | Edge AI inference SoC |
| d-Matrix | Series D | $250M | Digital in-memory compute inference chip |
| Etched | Series B | $150M | Transformer-specific ASIC |
| Tenstorrent | Series D | $100M | RISC-V AI accelerator |
Next Week's Watchlist
- NVIDIA GTC Taipei 2026 (June 10) — potential Rubin architecture details
- FERC review progress on AI data center grid connection applications
- TSMC June CoWoS-L capacity expansion target
- OpenAI / Oracle mega-cluster contract announcement
- EU AI Act energy efficiency requirements for large-scale training clusters
Disclaimer: Data and forecasts in this weekly report are based on public information and industry estimates. POC.HK strives for accuracy, but some forward-looking indicators are model-based estimates and may deviate from actual figures.