Nvidia’s 20-Year CUDA Moat Erodes as Open AI and Custom Chips Gain Ground
Updated
Updated · microwire.info · Aug 1
Nvidia’s 20-Year CUDA Moat Erodes as Open AI and Custom Chips Gain Ground
3 articles · Updated · microwire.info · Aug 1
Summary
Google TPUs, AWS Trainium and Inferentia, and in-house chips from Meta and Microsoft are reducing reliance on Nvidia by shifting AI workloads onto custom ASICs that bypass CUDA.
High-level frameworks such as PyTorch, Triton, and JAX further weaken Nvidia’s software lock-in because developers increasingly write tensor code above the hardware layer rather than directly for CUDA.
Western sanctions on top-end GPUs are portrayed as accelerating China’s self-sufficiency, pushing Huawei’s open-source CANN, domestic chipmakers, and DUV-based manufacturing toward a separate AI stack.
Chinese open-source models including Qwen, DeepSeek, Kimi, and Zhipu GLM are gaining global downloads and derivatives, while efficient designs such as MoE and MLA cut compute needs and support non-CUDA deployment.
The report argues the AI market is moving from a GPU monoculture to a split landscape of hyperscaler ASICs in the West and open-source silicon ecosystems in the East.