SemiAnalysis: Kimi K3's KDA mechanism improves attention efficiency but will require more GPU, HBM, DRAM, and network resources, not fewer!
Although Kimi K3's linear attention has triggered short-term concerns about hardware requirements, SemiAnalysis pointed out that its 2.8 trillion parameters and large-scale inference architecture will actually increase the demand for high-end GPUs, HBM, and high-speed interconnects.
On the dark side of the moon, the large-scale model Kimi K3 uses a linear attention mechanism, triggering concerns in the market that demand for NVIDIA, HBM, and network equipment may be weakened.
However, the semiconductor research firm SemiAnalysis recently provided a completely different assessment: the massive parameter scale and inference architecture requirements of K3 will not only not weaken the demand for high-end AI hardware, but may further strengthen the demand for NVIDIA's high-end GPUs, HBM, and high-speed interconnect devices.
SemiAnalysis pointed out that the parameter scale of K3 exceeds 2.8 trillion, with a model weight capacity exceeding 1.5TB HBM. Even in relatively limited user concurrency scenarios, a large amount of KV cache still needs to be offloaded to CPU DDR5 memory and NVMe storage, and there will not be a significant surplus of HBM space.
More importantly, it was revealed earlier that K3's efficient inference deployment requires a large-scale expansion domain architecture consisting of at least 64 chips. This hardware requirement is highly compatible with the design direction of NVIDIA's GB200/GB300 NVL72 and other rack-level AI systems.
SemiAnalysis believes that the market's previous understanding of linear attention as "weakening GPU demand" is based on a misconception. The real impact may be quite the opposite: more efficient model architectures reduce AI inference costs, which will drive more applications to be deployed, thereby stimulating long-term demand for GPUs, HBM, DRAM, and network infrastructure.
Unafraid of linear attention iteration, NVIDIA chip demand remains strong
Market concerns mainly stem from the Kimi Delta Attention (KDA) mechanism used by Kimi K3.
Compared to traditional Transformer attention mechanisms, KDA can significantly reduce the data transmission requirements of KV cache, reducing up to about 10 times the network bandwidth pressure. Some investors have been reminded of the market's concerns about AI hardware demand after the release of DeepSeek R1, believing that the improvement in model efficiency may reduce reliance on high-end computational hardware.
However, SemiAnalysis believes that this judgment overlooks another core demand in large-scale model inference the computational and interconnection pressure brought about by parameter scale.
K3 has over 2.8 trillion parameters, and its model weight itself requires deployment with a large-scale distributed computing system. At the same time, K3 uses the WideEP (Wide Expert Parallelism) optimization strategy, distributing 896 expert modules to multiple GPUs, allowing a single GPU to only bear part of the expert weights, thereby improving computational efficiency.
However, WideEP also poses new challenges: the frequent data exchange between experts requires more powerful network interconnection capabilities. SemiAnalysis points out that the copper backplane interconnection architecture used by GB200/GB300 NVL72 provides an intra-rack bandwidth 18 times that of traditional DGX B200 systems, making it ideal for large-scale expert parallel inference tasks.
In other words, the communication requirements saved by KDA's KV cache may be partially offset by the weight exchange requirements brought about by WideEP, and the overall pressure on AI infrastructure may not significantly decrease.
64-chip expansion domain is not exclusive to NVIDIA
However, there are differing opinions in the market on this issue. Informed source GDP (@bookwormengr) pointed out that the "64-chip expansion domain" mentioned by the dark side of the moon does not necessarily mean the NVIDIA NVL72 solution. Huawei's Ascend 950 SuperPod also uses a 64-chip configuration and has similar NVLink-like unified bus (UB) memory expansion capabilities.
In terms of architectural capabilities, the Ascend 950 SuperPod can support expansion across 16 racks to 1024 NPUs, making it equally competitive in meeting the demands of large-scale model inference. Therefore, the growth in hardware demand brought about by K3 does not necessarily mean that NVIDIA will be the sole beneficiary, but the trend of strengthened demand for high-end AI interconnect systems remains clear.
Jevons' Paradox: Efficiency improvement in AI may drive hardware demand growth
SemiAnalysis further invokes Jevons' Paradox to explain trends in AI infrastructure.
This theory holds that when a technology improves resource utilization efficiency and reduces unit costs, demand often does not decrease but may increase due to an expanded scope of applications. Applied to the field of AI, after the linear attention mechanism reduces inference costs, it may drive more enterprises to deploy AI applications, leading to a further expansion of global AI inference scale and ultimately driving growth in demand for GPUs, HBM, DRAM, and high-speed network devices.
However, GDP remains cautious about this. While recognizing the long-term logic of Jevons' Paradox, he points out that KDA has practical implications for optimizing state storage in long-context tasks. Even with a huge model weight scale, due to the use of 4-bit quantization and highly sparse design, the actual memory pressure may be lower than market intuition.
He believes that the real variable to watch is whether the world's largest AI companiesOpenAI, Anthropic, and Google DeepMindhave already or will adopt similar linear attention mechanisms like KDA, DeepSeek CSA/HCA, etc. If these top AI companies adopt such architectures on a large scale, the demand for memory and interconnect resources in long-context inference may significantly decrease, which will become an important factor affecting the future structure of AI hardware demand.
Shift in market focus: Can the growth in AI demand offset efficiency improvements in architecture?
Overall, the emergence of Kimi K3 does not simply point to a decline in AI hardware demand. For NVIDIA, the real competitive focus may shift from "single-card performance" to "system-level capabilities" including high-speed interconnect, rack-level expansion, and large-scale inference optimization.
As model scales continue to expand, even as attention mechanisms continuously optimize, AI infrastructure will still face pressure from parameter scale, expert parallelism, data exchange, and improvements in inference throughput.
In the future, the core issue the market needs to focus on is not whether linear attention reduces the consumption of individual resources, but whether the growth in demand brought about by the expansion of AI applications can continue to exceed the resource savings brought about by architecture efficiency improvements.
This article is reprinted from "Wall Street See News," GMTEight editor: Jiang Yuanhua.
Related Articles

SUNAC (01918) issues 285 million incentive shares.

CR HOLDINGS (01911) subscribed to a total principal amount of USD 12 million in trust deposits.

Short positions in US stocks soar to historic highs, the "wall of fear" in the bull market continues to rise.
SUNAC (01918) issues 285 million incentive shares.

CR HOLDINGS (01911) subscribed to a total principal amount of USD 12 million in trust deposits.

Short positions in US stocks soar to historic highs, the "wall of fear" in the bull market continues to rise.

RECOMMEND





