Nvidia Puts Groq 3 LPX Into Full Production as AI Inference Demand Grows
Nvidia has moved its Groq 3 LPX rack into full production, marking a major step in the commercial rollout of technology obtained through its acquisition of AI chip startup Groq. The Groq 3 LPX...
Nvidia has moved its Groq 3 LPX rack into full production, marking a major step in the commercial rollout of technology obtained through its acquisition of AI chip startup Groq.
The Groq 3 LPX systems are expected to be deployed alongside Nvidia’s Vera central processors and Rubin graphics processors at neocloud provider Nebius, with initial deployments scheduled to begin later this year.
The move highlights Nvidia’s growing focus on low-latency AI inference, an increasingly important requirement as AI agents become capable of performing more complex tasks. Faster inference can reduce the delay users experience when interacting with AI systems, particularly in applications such as coding and other real-time workloads.
Nvidia’s $20 billion Groq acquisition
Nvidia acquired assets from Groq for approximately $20 billion in December, making it one of the company’s largest acquisitions.
Groq’s chip architecture is designed around high-speed inference and includes 500MB of high-speed SRAM directly on the chip. This architecture is intended to reduce memory-related bottlenecks when AI models generate responses.
Groq chips are manufactured by Samsung, while Nvidia’s GPUs are produced by Taiwan Semiconductor Manufacturing Company (TSMC).
Nvidia is packaging 256 Groq 3 chips into each LPX rack. According to Nvidia, a Groq 3 LPX rack can reach speeds of up to 3,400 tokens per second, based on a benchmark conducted by Artificial Analysis.
Groq targets low-latency inference
Nvidia’s move puts it in competition with other chipmakers and AI companies developing specialised hardware for fast inference.
AMD announced earlier this year that it planned to integrate its rack-scale systems with chips from Cerebras. OpenAI has also introduced an Ultrafast mode powered by Cerebras, with claimed speeds of up to 750 tokens per second.
However, Groq chips are not intended to replace Nvidia’s GPUs. GPUs remain the primary hardware used for demanding AI workloads, including model training and inference.
Instead, Groq’s technology focuses heavily on the “decode” stage of AI inference, where generated tokens are produced and delivered to users.
Nvidia expands Vera Rubin production
Nvidia is also increasing shipments of its Vera Rubin systems, which entered production earlier this year.
At the March unveiling of the Vera Rubin and Groq 3 LPX systems, Nvidia CEO Jensen Huang projected that Blackwell and Vera Rubin systems could generate a combined $1 trillion in cumulative sales through 2027.
With Groq 3 LPX now entering full production, Nvidia is expanding its AI infrastructure portfolio beyond traditional GPUs and positioning specialised low-latency hardware as an important part of the next generation of AI computing.





