OpenAI’s Jalapeño chip leads Blackwell in inference tests
At Hot Chips, OpenAI put its new Jalapeño system through its first public performance test. On Semianalysis’s InferenceX benchmark, it registered more tokens per user and more throughput per kilowatt than an Nvidia Blackwell system, according to TechCrunch. Richard Ho, OpenAI’s head of hardware, presented the result as a way to serve more AI work while returning answers faster.
The chip’s approach is to reduce the traffic around inference—the processing involved in generating a model’s response. OpenAI says Jalapeño keeps model state, including the KV cache used during generation, local to the resources that need it, then activates the appropriate mix of computing, memory and networking for each phase. That design targets delays in prefill, which prepares the request, and in communication between system components.
The hardware has a listed 700W TDP, or thermal design power. OpenAI developed it with Broadcom and plans to make Jalapeño a multigenerational platform, coordinating its AI products, models, chips and memory rather than treating the processor as an isolated component.
The timing matters. Ho estimated that Jalapeño would deploy in very small volumes at the end of 2026, with more significant deployment in 2027. The benchmark comparison is with an Nvidia Blackwell system available today, while other processors may have advanced by the time Jalapeño is deployed at scale.
So what changes in practice? If the benchmark advantage holds outside the test, AI operators could handle more users for each unit of power and reduce response delays. For now, the evidence is a benchmark result and a planned deployment—not a broad production rollout.
Comentários
A carregar a conversa…
Inicie sessão para escrever um comentário. Iniciar sessão