Google is currently facing a significant infrastructure challenge as demand for artificial intelligence inference outpaces the company’s ability to supply computing power. According to a report by the Financial Times, the tech giant was unable to fulfill the full Gemini inference capacity requests made by Meta Platforms earlier this year.
The shortfall underscores a shift in the AI industry: while the initial focus was on training large models, the current bottleneck is now centered on inference—the compute power required to execute tasks in real-time. Unlike model training, which occurs periodically, inference happens billions of times daily, creating massive strain on data centers.
Alphabet CEO Sundar Pichai has acknowledged that Google Cloud revenue could have been higher if the company possessed more available capacity. Despite Alphabet investing over $90 billion in 2025 and planning to double that figure this year to expand its custom Tensor Processing Units (TPUs) and data center footprint, supply remains scarce.
The inability to meet demand is not limited to Meta. Industry observers note that Google continues to manage access for various enterprise clients as it scales its hardware infrastructure. Analysts suggest this supply-side constraint indicates that the AI boom is currently limited by physical hardware production—including GPUs, memory, and power systems—rather than a lack of market interest.

