Morgan Stanley: AI Memory Shortage to Persist for Years, Three Paths to Break the Deadlock, Structural Opportunities in Storage and Heterogeneous Computing
Morgan Stanley released a semiconductor industry research report stating that the DRAM memory shortage will persist throughout this AI industry cycle, and AI computing power expansion will not wait for new wafer fabs to be built and put into production.
Morgan Stanley released a semiconductor industry research report noting that the DRAM memory shortage will persist throughout this AI industry cycle, and AI compute expansion will not wait for new wafer fabs to be built and come online. The industry is bypassing the memory bottleneck through three major technological pathshardware de-speccing, inference architecture disaggregation, and CXL memory poolingto continue advancing compute buildout under supply constraints. The bank remains bullish on memory leaders such as Micron (MU.US) and SanDisk (SNDK.US), while also favoring incremental opportunities in the CXL interconnect and heterogeneous inference tracks, recommending Astera Labs (ALAB.US), Marvell (MRVL.US), Cerebras (CBRS.US), and NVIDIA Corporation (NVDA.US).
Morgan Stanley emphasized that the current AI memory shortage is not a short-term disruption, but a structural contradiction in which compute performance growth far outpaces memory supply growth. Frontier large model scale doubles every 6 months, leading models' context windows are growing at 5-6x annually, and combined with continuously rising inference concurrency demand, memory capacity and bandwidth requirements continue to surge. Meanwhile, DRAM fab construction cycles last for years, supply elasticity is extremely low, and the shortage pattern will persist for years. NVIDIA Corporation CEO Jensen Huang has also publicly stated that the industry needs to shift its thinking and address memory constraints through architectural innovation rather than simply waiting for capacity expansion.
Facing tight HBM and main memory supply, the industry's most direct response is selectively reducing per-device memory specifications (De-speccing). Taking NVIDIA Corporation's Rubin architecture as an example, per-rack LPDDR5 capacity was reduced from the originally planned 54TB to 28TB, and per-GPU HBM capacity was lowered from 288GB to 192GB, ensuring overall system shipment volumes by reducing memory stack height. Morgan Stanley noted that de-speccing does not eliminate the memory bottleneck but shifts pressure across memory tiers: after cutting high-speed local memory, non-high-frequency data such as KV cache migrates down to NAND storage, while cross-GPU data interaction increases, driving demand for network interconnect bandwidthessentially supplementing scarce high-bandwidth memory with faster network interconnect and lower-cost storage resources.
The second path to breaking the deadlock is inference architecture disaggregation. AI inference comprises two significantly different stages: Prefill and Decode. Prefill is primarily compute-intensive, while Decode is highly dependent on memory bandwidth. Traditional architectures use the same accelerator for both types of tasks, resulting in relatively low resource utilization. The industry is now accelerating toward heterogeneous inference, splitting the two stages onto different hardware: compute-intensive Prefill is handled by general-purpose GPUs, while memory-bandwidth-intensive Decode is processed by specialized accelerator chips equipped with large on-chip SRAM.
Typical examples include Cerebras's wafer-scale engine and the Groq LPU architecture acquired by NVIDIA Corporation, both of which can deliver memory bandwidth efficiency far exceeding traditional GPUs during the decode stage. Morgan Stanley believes that heterogeneous inference will become an important evolutionary direction for AI infrastructure, and compute vendors specializing in the decode segment will gain a clear incremental market.
The third path is CXL technology restructuring the memory system. CXL (Compute Express Link) enables memory expansion, sharing, and pooling through high-speed interconnect, decoupling memory from a single processor and becoming a core technological path to breaking through memory capacity constraints. Morgan Stanley estimates that AI demand will drive the CXL and related memory attach chip market to approximately $6 billion by 2030, far exceeding the traditional CPU memory expansion market size. Its core value lies in building a tiered memory architecture: the highest-frequency data remains in HBM, sub-high-frequency data is placed in CXL memory pools, and low-frequency cold data sinks to NAND, matching data access frequency with cost gradients.
Regarding investment themes, Morgan Stanley reiterated three major directions: first, continue to overweight Micron and SanDiskde-speccing stems from supply shortages rather than weakening demand, AI memory demand is trending upward long-term, and the supply gap will continue to absorb capacity; second, position in CXL and scale-up interconnect leaders Astera Labs and Marvell; third, focus on heterogeneous inference beneficiary Cerebras, as well as NVIDIA Corporation, which is completing its heterogeneous layout through the acquisition of Groq.
Risk warnings: AI compute demand growth falling short of expectations; CXL technology adoption progressing slower than expected; memory capacity expansion exceeding expectations.
Related Articles

AI compute demand outlook viewed favorably; BNP Paribas raises target prices for NVIDIA Corporation (NVDA.US) and Intel Corporation (INTC.US), but remains cautious on the latter.

SINOPHARM TECH (08156) restores public float

QINGLING MOTORS (01122): Enters into Xinjiang Lixin Energy commercial vehicle repurchase agreement, maximum repurchase price not exceeding RMB 12.636 million
AI compute demand outlook viewed favorably; BNP Paribas raises target prices for NVIDIA Corporation (NVDA.US) and Intel Corporation (INTC.US), but remains cautious on the latter.

SINOPHARM TECH (08156) restores public float

QINGLING MOTORS (01122): Enters into Xinjiang Lixin Energy commercial vehicle repurchase agreement, maximum repurchase price not exceeding RMB 12.636 million






