Thursday, 27 August 2026

Aube.

News of progress
In productionCorroborated · 3 sources

NVIDIA puts Groq 3 LPX into production for AI inference

Languages for this article
Original · ENFR

Originally written in English. 2 languages available; yours is one click away.

An AI agent checking files, writing and testing code, calling tools and validating its output has one recurring bottleneck: generating the next pieces of text quickly enough to keep the workflow moving. NVIDIA says Groq 3 LPX, an accelerator for AI inference, is now in full production as part of its Vera Rubin platform, aimed at increasing token-generation speed for these agentic systems.

The company reported 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context window, citing benchmark results from Artificial Analysis. NVIDIA presents that figure as a reference point for high-context agentic inference, and says Groq 3 LPX can reduce completion time for some coding workflows and improve responsiveness in latency-sensitive uses.

The logic is straightforward: agentic systems often generate many tokens across many inference steps, while also working with large amounts of context. Groq 3 LPX is designed to raise the generation rate for individual users. Vera Rubin NVL72 systems are intended to support both training and inference, with the LPX extending inference performance in rack-scale AI factory deployments.

Nebius's cloud deployment is still a plan. Nebius intends to deploy Groq 3 LPX in Nebius Token Factory, its production inference platform, to provide higher token-generation throughput for developers building agentic AI applications. Groq also plans to be an early adopter. NVIDIA describes the wider Vera Rubin platform as spanning seven chips and five rack types, including BlueField-4 networking and storage components.

So what does this change in practice? If the planned deployments deliver the stated performance, developers could get more responsive agents for workloads that repeatedly call tools, inspect data and test code, especially when context windows are large. The boundary is equally clear: the 3,400-token result is a benchmark NVIDIA reported from Artificial Analysis, and the product’s benefit will vary by model, context and workload.

3,400 output tokens per secondReported generation rate on Gemma 4 31B with a 100,000-token context

Sources — read the originals(Paris time)

The Korea Herald — TechEN
Data Center DynamicsEN
Engineering.comEN
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in