xAI is building Colossus 2, a gigawatt-scale training datacenter, leveraging innovative power sourcing across state lines and partnerships with Solaris Energy. The project aims to surpass rivals by Q3 2025, with potential Middle East funding.
Nvidia announced the Rubin CPX, a solution specifically optimized for the prefill phase, emphasizing compute FLOPS over memory bandwidth. This is a game changer for inference, second only to the March 2024 announcement of the GB200 NVL72 Oberon rack-scale form factor. Specialized hardware for prefill and decode unlocks the full potential of disaggregated serving. Nvidia's rack system design gap has become canyon-sized, forcing competitors to reconfigure their roadmaps.
Huawei is ramping up Ascend AI chip production using a die bank from TSMC and capacity from SMIC. However, HBM (High Bandwidth Memory) shortages will become the primary bottleneck for future production. China's domestic HBM supplier CXMT is ramping quickly but cannot meet demand in the near term. The article also analyzes the impact of export controls and the potential implications of Nvidia H20 chip sales to China.
Two-and-a-half years ago, SemiAnalysis flagged a looming 'cloud crisis' at AWS. Today, evidence mounts as Azure leads quarterly cloud revenue and Google Cloud narrows the gap. Yet SemiAnalysis makes an out-of-consensus call for an AWS AI resurgence, driven by its partnership with Anthropic. Anthropic's revenue surged from $1B to $5B annualized in 2025, and AWS is building over 1.3GW of datacenter capacity for it, hosting nearly a million Trainium2 chips. Despite Trainium2 lagging Nvidia on specs, its memory bandwidth per TCO advantage aligns with Anthropic's reinforcement learning roadmap. The collaboration is evolving into a custom silicon program, poised to boost AWS growth above 20% by end of 2025.
This article provides an in-depth analysis of H100 and GB200 NVL72 training benchmarks, covering model flops utilization (MFU), total cost of ownership (TCO), cost per million tokens, energy consumption, and reliability. It reveals that H100 achieved up to 57% throughput improvement over 12 months via software optimization alone. Meanwhile, GB200 NVL72 offers potential performance advantages but faces reliability challenges and has not yet completed large-scale training runs. Detailed benchmarks for models like GPT-3 175B and Llama 3 405B are presented, along with three recommendations for Nvidia: increase benchmark transparency, expand to native PyTorch, and improve GB200 diagnostic tools.
GPT-5's release disappointed power users, but the real focus is on monetizing ChatGPT's 700M+ free users. The article analyzes how OpenAI's new router technology distinguishes query intent and enables future monetization through agentic purchasing and transaction fees, potentially creating a consumer superapp.
Robots have powered manufacturing for decades, yet they stayed single-purpose and thrived only in perfect settings. Previous attempts at intelligent machines overpromised and underdelivered. But they were too early. Today, modern AI paradigms convert most robot roadblocks into data problems and push machines toward capabilities once thought impossible. As these models absorb real-world experience, robots will sharpen current skills, gain new ones, and deploy faster, absorbing ever-increasing shares of labor.
Meta’s shocking purchase of 49% of Scale AI at a ~$30B valuation shows that money is of no concern for the $100B annual cashflow ad machine. Despite seemingly unlimited resources, Meta has been falling behind foundation labs in model performance.