In the era of AI reasoning, the demand for TPUs is expanding explosively! Morgan Stanley: Google's million TPUs leverage $200 billion in cloud computing revenue.

date
22:01 25/08/2026
avatar
GMT Eight
According to institutions like Morgan Stanley and Wedbush Securities, which remain optimistic about the investment prospects of the AI computing power industry chain, the nearly limitless demand for cutting-edge computing power in the era of AI inference, alongside the computational needs surrounding AI agents, allows AI ASICs to grow into an important component of a second trillion-dollar computing ecosystem without undermining the demand for GPUs.
Morgan Stanley, the Wall Street financial giant, has recently estimated that as the AI computing power industry moves from centralized AI training to large-scale AI inference, the AI ASIC computing power clusters led by Alphabet Inc. Class C TPU are transitioning from internal cost-reduction and customization tools within Alphabet Inc. Class C's cloud computing business to highly profitable AI cloud infrastructure services available for external sales. With a commitment for approximately $35 billion for one million TPUs, Morgan Stanleys analyst team has significantly raised its price assumption for Alphabet Inc. Class C TPU systems from $20 billion per gigawatt to $27 billion and increased its gross margin assumption from 20% to 30%. In a recent research report, Morgan Stanley projects that Alphabet Inc. Class C's cloud revenue related to TPU will reach $84 billion and $108 billion in 2027 and 2028, respectively, totaling nearly $200 billion. This suggests that Alphabet Inc. Class C's parent company, Alphabet (GOOGL.US), could further transform from a major player in AI capital expenditures to a super platform for renting/selling AI computing power systems, with TPU potentially becoming another significant revenue growth engine after search ads and public cloud computing IaaS + SaaS businesses. One million TPUs unlock a $35 billion order! Morgan Stanley: Alphabet Inc. Class C aims at a $200 billion computing power gold mine. For mature models with relatively stable structures, heavy usage, and long-term high utilization rates, along with open-source models and AI agent workloads, dedicated AI chips like TPU (the typical AI ASIC technology developed by major cloud computing firms) can be deeply optimized around low-precision matrix computation, memory access, and chip interconnects, thus improving costs per Token, throughput per watt, and total cost of ownership. At the same time, Alphabet Inc. Class C TPU computing clusters can also handle large-scale training tasks, indicating a more nuanced trend of heterogeneous computing power division between GPUs and ASIC/TPU. This means that the AI training operator processes and the most complex and rapidly changing cutting-edge AI workloads still heavily depend on GPU clustersfrontier model pre-training, reinforcement learning, and rapidly evolving new operators continue to rely on the programmability of GPUs, the CUDA ecosystem, and NVLink/NVSwitch cluster capabilities. Meanwhile, large-scale AI inference workloads surrounding mature and open-source AI large models, Copilot, and AI Agents (AI intelligent agents) are increasingly well-suited for dedicated self-developed AI chips. Specifically, Alphabet Inc. Class C has divided its eighth-generation TPU into training type TPU 8t and inference type TPU 8i: the 8i features larger on-chip SRAM, 288GB HBM, and a dedicated aggregate communication engine that optimizes KV Cache and low-latency inference, claiming its inference cost-performance has improved by as much as 80% compared to Ironwood. At institutions like Morgan Stanley and Wedbush Securities, which maintain a positive outlook on AI computing power investments, the seemingly endless demand for cutting-edge computing power in the AI inference era, alongside the computing needs surrounding AI agents, enables AI ASICs to grow into a critical component of a second trillion-dollar computing power ecosystem without destroying GPU demandfurther reinforcing the investment logic in the AI computing power industry chain that the AI semiconductor super cycle is not a single GPU cycle but a period of exponential improvement in the silicon content of entire data centers. Morgan Stanley analyst Brian Novak's team wrote in their latest report to clients, It is worth noting that recent media reports indicate that Alphabet's Alphabet Inc. Class C has committed to supplying about one million TPUs for approximately $35 billionour estimates indicate that this equates to about 1.3 gigawattsessentially corresponding to our revised assumption of approximately $27 billion per gigawatt. Previously, we assumed Alphabet Inc. Class C would sell racks and systems to these first-party partnerships at a price of $20 billion per gigawatt, with a gross margin of 20%. Furthermore, the expanded customized chip development agreement between Alphabet Inc. Class C and Marvell Technology, Inc.which our semiconductor research team has previously analyzedaligns with our assessment that TPU pricing and average selling prices exceed our prior assumptions, and can be seen as supportive evidence for this judgment. In light of this new information, we have significantly raised our expectations for first-party TPU sales to $27 billion per gigawattwith the gross profit margin rising sharply to 30%. Overall, we currently project that Alphabet Inc. Class C will sell 0.3 gigawatts of TPU systems in the second half of 2026, 3.2 gigawatts in 2027, and 4.2 gigawatts in 2028. Based on these calculations, the TPU-related revenue figures for Alphabet Inc. Class C's cloud computing business are expected to reach $84 billion and $108 billion in 2027 and 2028, respectively, wrote Brian Novak's analyst team. The analyst team led by Novak continues to give Alphabet Inc. Class C an overweight rating, maintaining a target price of $400. As of Tuesdays early trading, Alphabet's stock price hovered around $348 per share. Alphabet Inc. Class C collaborates with Marvell, and Microsoft Corporation's Maia 200 is on the verge of launch, while AI giants compete for silicon supremacy in the AI inference era. The Token economics of the AI inference era are fundamentally different from those of AI training workloads, which include forward propagation, backpropagation, gradient synchronization, and frequently changing research-driven operators, placing greater emphasis on generality, software maturity, and cluster scalability; inference can instead be divided into computation-intensive pre-filling and memory bandwidth, KV caching, and delay-sensitive decoding. ASICs/XPUs can eliminate unnecessary general-purpose circuits and harden for FP8/FP4 low-precision matrix operations, attention mechanisms, mixture of experts models (MoE), KV caches, and data movement paths, thereby enhancing the number of Tokens per watt, dollars per Token, and the determinism of service-level agreements (SLAs). It is important to emphasize that while the model invocation patterns of AI agents are more suited to ASICs than GPUs, the orchestration, tool invocation, retrieval augmentation, and state management still require collaboration among CPUs, memory, networks, and storage, thereby further reinforcing heterogeneous computing systems in the era of agents. One of Americas tech giants, Microsoft Corporation (MSFT.US), plans to soon introduce its latest self-developed AI accelerator, Maia, alongside the new agreement between Alphabet Inc. Class C and fabless chip company Marvell, covering custom chips such as AI accelerators, storage, and memory controllers. This highlights the process of AI computing power infrastructure construction led by AI application pioneers and hyperscale cloud service providers, which is undergoing a shift from collective procurement of NVIDIA Corporation GPUs to a large-scale collaborative operation of heterogeneous computing systems that includes NVIDIA Corporation GPUs, AMD GPUs, self-developed AI-specific integrated circuits (AI ASICs)/XPUs, and CPUs/DPUs. As the architecture of large AI models stabilizes and the global Token invocation volume shows exponential growth, the high concurrency AI inference workloads will increasingly benefit from customized chips that reduce unit Token costs. The new agreement between Alphabet Inc. Class C and Marvell (Marvell Technology, Inc.) comprehensively covers AI accelerators, storage, and memory controllers, with potential procurement corresponding to an estimated maximum revenue of about $120 billion for Marvell through fiscal year 2033. This effectively diversifies the long-standing reliance on Broadcom Inc., Alphabet Inc. Class C's sole technology partner for TPU. Microsoft Corporations Maia 200, with Taiwan Semiconductor Manufacturing Co., Ltd.s 3nm technology, FP8/FP4 tensor cores, 272MB on-chip SRAM, and custom on-chip networks (NoC), compresses costs for Azure, Copilot, and large model inference. Its relative advantage over NVIDIA Corporation GPUs is not that any model is faster, but rather that Microsoft Corporation can jointly optimize chips, compilers, models, data center networks, and cloud scheduling, avoiding paying all profit premiums to external GPU vendors. A research report from Microsoft Corporation claims that Maia 200 has achieved a 30% increase in performance per dollar compared to the latest generation hardware in its fleet, and that when running its self-developed MAI model on Maia 200, performance per watt has increased by about 40%. This is the most dangerous competitive edge of ASICs/XPUs in the inference era: once billions of similar Token generation tasks can be highly standardized, the flexibility premium of NVIDIA Corporations universal GPU may not justify payment for every single inference request.