Nvidia said its Groq 3 LPX rack has entered full production, turning technology from its $20 billion Groq acquisition into a shipping product. The racks will deploy at neocloud Nebius alongside Vera CPUs and Rubin GPUs and go online later this year, and Nvidia is positioning them for the low-latency inference that AI coding agents need.
Nvidia's Groq 3 LPX rack is now in full production, the company said Monday, marking the commercial debut of technology from its largest acquisition on record. The rack will run alongside Vera central processors and Rubin graphics processors at neocloud Nebius and will be online later this year, Nvidia senior director Dion Harris told reporters.
Low-latency chips target responsive AI agents
The push to ship Groq's chip highlights the growing importance of low-latency inference, which keeps AI agents from lagging when users interact with them, especially for coding tasks. Cloud providers can charge more for this kind of low-latency token, according to Nvidia. Harris said the technology unlocks the ability to offer premium service tiers for customers who need the fastest response times.
The $20 billion Groq deal comes to market
Nvidia bought assets from chip startup Groq in December for $20 billion, its largest purchase to date. The Groq architecture packs 500 megabytes of fast SRAM directly onto the chip's die to cut memory bottlenecks, and Samsung manufactures the chips, while Taiwan Semiconductor Manufacturing Co. builds Nvidia's GPUs. Nvidia packs 256 individual Groq 3 chips into each LPX rack, and it says the rack can deliver 3,400 tokens per second, citing a benchmark from Artificial Analysis.
It's a competitive space. Advanced Micro Devices said earlier this year it would integrate its rack-scale systems with chips from Cerebras, which recently went public, and OpenAI's newly announced Ultrafast mode currently promises 750 tokens per second running on Cerebras hardware.
GPUs still do the heavy lifting
Low-latency chips do not replace the GPU, which remains capable of both training and inference and stays flexible across new models. Chips like Groq instead focus on the "decode" phase of serving a model. According to CNBC: "This isn't about replacing GPUs", Harris said, adding that it's about matching the right processor to the right part of the workload.
Nvidia is also ramping shipments of its Vera Rubin systems, which entered production earlier this year. At the March unveiling of Vera Rubin and Groq 3 LPX, CEO Jensen Huang projected $1 trillion in combined sales from current-generation Blackwell chips and the new Vera Rubin systems through 2027.
Huang said he would allocate a quarter of data center space earmarked for coding applications to Groq chips, with the rest running on Vera Rubin. Nvidia is scheduled to report earnings on Wednesday.
Source: CNBC
Trading involves risk.