Updated
Updated · CNBC · Aug 24
Nvidia Puts 256-Chip Groq 3 LPX Racks Into Production as $20 Billion Deal Reaches Market
Updated
Updated · CNBC · Aug 24

Nvidia Puts 256-Chip Groq 3 LPX Racks Into Production as $20 Billion Deal Reaches Market

3 articles · Updated · CNBC · Aug 24

Summary

  • Groq 3 LPX racks are now in full production, with Nvidia saying the systems will go online at neocloud Nebius later this year alongside Vera CPUs and Rubin GPUs.
  • The rack packages 256 Groq 3 chips and delivers 3,400 tokens per second, targeting low-latency inference for AI agents and coding workloads where cloud providers can charge premium rates.
  • Nvidia bought Groq assets for $20 billion in December, and the launch marks the first commercialization of that acquisition; the chips use 500 megabytes of on-die SRAM and are made by Samsung.
  • The push comes as AMD teams with Cerebras on rack-scale inference systems and OpenAI's Cerebras-powered Ultrafast mode advertises 750 tokens per second, underscoring a crowded market for faster AI serving.
  • Nvidia says Groq chips are meant to handle the decode phase rather than replace GPUs, fitting into a broader Vera Rubin rollout that CEO Jensen Huang has tied to $1 trillion in cumulative Blackwell-and-Rubin sales through 2027.

Insights

Will Nvidia's $20 billion gamble on Groq reshape AI economics, or is hardware specialization a costly misstep?
As AI inference splits into phases, who will truly control the lucrative market for ultra-fast token generation?