SemiAnalysis: K3's linear attention is bullish, not bearish, for NVIDIA/HBM/networking
SemiAnalysis: K3's linear attention is bullish, not bearish, for NVIDIA/HBM/networking
8-part technical rebuttal of the K3-kills-compute panic: 2.8T+ params need rack-scale NVL72 scale-up domains; WideEP (896 experts) trades lower KV-cache networking for MORE weight-shuffling bandwidth; >1.5TB of HBM for weights pushes KV-cache to DDR5/NVMe; Kimi says optimal serving needs 64+ chips; Jevons closes it. The hardware-demand half of the desk's K3 call, argued from the architecture.
Why it's worth your time
8-part technical rebuttal of the K3-kills-compute panic: 2.8T+ params need rack-scale NVL72 scale-up domains; WideEP (896 experts) trades lower KV-cache networking for MORE weight-shuffling bandwidth; >1.5TB of HBM for weights pushes KV-cache to DDR5/NVMe; Kimi says optimal serving needs 64+ chips; Jevons closes it. The hardware-demand half of the desk's K3 call, argued from the architecture.
Full thread (posted 2026-07-17 3:49 PM, ~459K views at capture)
Similar to the panic over DeepSeek R1, some uneducated people think Kimi K3's use of linear attention (KDA) is bad for NVIDIA, HBM, DRAM, and networking because it has relatively lower KV-cache requirements. The opposite is true, and we explain why below. 1/8
Kimi K3 is actually quite positive for NVIDIA, as large-model inference is where the NVL72 shines. Because K3 has more than 2.8 trillion parameters, it requires a large scale-up domain to store its weights. 2/8
Secondly, although Kimi Delta Attention has up to 10× lower networking requirements for KV-cache transfers, its large weights require even more network bandwidth to implement an optimization called WideEP, which spreads the weights across different GPUs. 3/8
WideEP distributes the 896 experts across many GPUs so that each GPU's HBM contains only a small number of experts, optimizing per-token memory usage and compute utilization. 4/8
The unfortunate downside of the WideEP optimization is that it consumes a tremendous amount of network bandwidth. WideEP is highly optimized for rack-scale systems like the GB200/GB300 NVL72, whose copper backplane provides 18× more bandwidth than comparable DGX B200 systems. 5/8
Furthermore, since the weights occupy more than 1.5 TB of HBM capacity, the KV cache for K3's KDA and Gated MLA will need to be offloaded to CPU DDR5 and NVMe, even at relatively low user concurrency, because little space remains in HBM. 6/8
Kimi themselves have stated that optimal K3 inferencing will require a rack with a large scale-up domain with at least 64 chips. 7/8
Lastly, Jevons' Paradox means that making attention more efficient will drive wider AI adoption, which will ultimately require more GPUs, HBM, DRAM, and networking—not less. 8/8
Related
8 eventsSource links
Tweets, videos, filings, articles, and source pages tied to this read.