Technology

Groq to Deploy NVIDIA Groq 3 LPX for AI Inference Cloud

Groq will be among the first to use NVIDIA's Groq 3 LPX accelerators with the Vera Rubin NVL72 platform in its AI inference cloud.

By Tim Editorial

Groq to Deploy NVIDIA Groq 3 LPX for AI Inference Cloud
nvidia.com

Groq has announced that it will be one of the first adopters of NVIDIA's Groq 3 LPX, an AI inference accelerator designed for agentic systems. The announcement was made via Groq's official account on X on Monday, August 24, 2026. In its statement, Groq said it will use the Groq 3 LPX alongside NVIDIA's Vera Rubin NVL72 platform in its AI inference cloud infrastructure, which is being built specifically for this purpose. Groq noted that it is working with Dell Technologies to deploy the NVIDIA Groq 3 LPX. This collaboration is part of Groq's strategy to deliver ready to use inference computing capacity for enterprises and AI developers building next generation agents.

According to Groq, once the NVIDIA Groq 3 LPX capacity becomes operational, the infrastructure will already be optimized to handle high demand inference workloads. The NVIDIA Groq 3 LPX is an interactive inference accelerator specifically designed for agentic AI systems. Based on information from NVIDIA's official website, the product is positioned as an AI inference solution for agentic systems that require low latency and large context. NVIDIA describes the Groq 3 LPX as a rack scale inference accelerator developed for the NVIDIA Vera Rubin platform. In terms of technical specifications, a StorageReview report published on March 19, 2026, stated that the NVIDIA Groq 3 LPX combines 256 LP30 LPUs with the Vera Rubin NVL72 platform.

This combination is said to deliver up to 35 times higher inference throughput per megawatt compared to the previous generation. This figure is one of the main attractions for cloud providers looking to improve inference computing efficiency without significantly increasing power consumption. NVIDIA explained in its technical blog that the Groq 3 LPX was designed together with the NVIDIA Vera Rubin platform to meet the low latency and large context requirements characteristic of agentic systems. This architecture allows AI models to process large numbers of requests with faster response times, a need that is becoming more urgent as generative AI applications and autonomous agents expand. For Groq, this move is part of its efforts to strengthen its position in the AI inference infrastructure market.

Groq has been known as the developer of the LPU (Language Processing Unit), a processor architecture specifically designed to run large language models with low latency. By adopting the NVIDIA Groq 3 LPX, Groq adds NVIDIA based computing capacity alongside its own LPU technology. This announcement also underscores the intensifying competition in the AI accelerator market. NVIDIA continues to expand its product line from GPUs for model training to specialized solutions for inference. The Groq 3 LPX is a response to the market's need for more efficient and responsive inference infrastructure, particularly for agentic applications that require real time user interaction. Dell Technologies, as the deployment partner, is also a key part of this ecosystem.

Dell's involvement shows that the adoption of the NVIDIA Groq 3 LPX involves not only cloud providers and chipmakers but also infrastructure vendors that supply server hardware and supporting systems. Such collaborations are becoming a common pattern in the AI industry, where companies with different expertise join forces to deliver complete computing solutions. Groq stated that enterprises and AI developers building next generation agents will get one of the earliest paths to use the NVIDIA Groq 3 LPX on real production workloads. This statement underscores that Groq is targeting users who need large scale inference capacity ready for commercial applications, not just trials or research.

From a market perspective, Groq's move reflects a broader trend in the AI computing industry: a shift in focus from model training to inference. As more AI models mature and become ready for deployment, the need for fast, efficient, and power efficient inference infrastructure becomes a top priority for cloud service providers and technology companies. Products like the NVIDIA Groq 3 LPX are designed to address this specific need. There is no official information yet on the exact timeline for when the NVIDIA Groq 3 LPX capacity will become operational in Groq's infrastructure. However, this announcement signals that the necessary technical preparations and partnerships are already underway.

Groq has also not disclosed details on the computing capacity it will provide or the scale of investment for this project. Going forward, Groq's success in bringing the NVIDIA Groq 3 LPX to market will depend on its technical execution and deployment speed. Competition in the AI inference cloud provider market is intensifying, with major players like AWS, Microsoft Azure, and Google Cloud also continuously updating their infrastructure. Groq needs to ensure that its services truly deliver the latency and efficiency advantages promised by the Groq 3 LPX architecture. For developers and companies building agentic AI applications, the availability of low latency inference infrastructure is crucial.

Applications such as virtual assistants, automated customer service agents, and real time decision making systems require very fast response times to provide a good user experience. The NVIDIA Groq 3 LPX, with its combination of 256 LPUs and the Vera Rubin NVL72 platform, is designed to meet these needs. Groq's announcement also indicates that the AI hardware ecosystem is evolving rapidly. NVIDIA is not only competing in the GPU segment for training but also building a dedicated product line for inference with better power efficiency. Meanwhile, companies like Groq, which originally developed their own processor architectures, are now choosing to also adopt NVIDIA technology, showing that the market is large enough for different architectural approaches.

Sources and references