OpenAI’s Jalapeño AI Chip Targets Nvidia With Faster, More Efficient AI Inference
OpenAI is expanding its focus beyond AI models and into custom hardware with Jalapeño, its first in-house inference chip developed in collaboration with Broadcom. The company says the chip is...
OpenAI is expanding its focus beyond AI models and into custom hardware with Jalapeño, its first in-house inference chip developed in collaboration with Broadcom. The company says the chip is designed to deliver faster AI responses while using power more efficiently.
Jalapeño is built specifically for AI inference, meaning it is designed to run trained models and generate responses rather than train new models. OpenAI has tested the chip with GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T to evaluate its performance across different workloads.
OpenAI’s Jalapeño chip performance
According to OpenAI’s published benchmark results, Jalapeño delivered between 1.5 and 1.9 times more AI work per watt than the comparison systems across the three tested models. The company also reported 1.7 to 3.6 times lower end-to-end latency. For highly interactive workloads, OpenAI says Jalapeño achieved between 2.1 and 4.1 times higher performance.
The benchmarks were conducted using InferenceX, a public benchmark developed by SemiAnalysis. OpenAI compared Jalapeño against Nvidia systems including the GB200 and GB300, depending on the model being tested.
For GPT-OSS 120B, OpenAI reported around 1.9 times higher peak throughput per watt compared with the GB200. On DeepSeek R1 670B, Jalapeño delivered about 1.7 times the throughput per watt of the GB300, while Kimi K2.5 1T showed an advantage of roughly 1.5 times.
Jalapeño focuses on speed and efficiency
One of the main goals behind Jalapeño is to address the traditional trade-off between throughput and latency. OpenAI says its architecture is designed to provide both higher throughput and lower response times, which could be particularly useful for AI agents that need to complete multiple steps quickly.
The chip is rated at 700 watts, although OpenAI says its measured sustained power remained at or below 550 watts during the workloads tested. OpenAI plans to deploy Jalapeño in large-scale systems, with a 128-chip configuration capable of 1.7 exaflops of 4-bit compute and 27.5TB of HBM4 memory.
What this means for Nvidia
Jalapeño represents a broader effort by major AI companies to gain greater control over the hardware powering their models. Custom chips can allow companies to optimise hardware around their own workloads while potentially improving efficiency and reducing dependence on third-party accelerators.
However, OpenAI is not abandoning Nvidia. The company says it will continue deploying accelerators from Nvidia and other hardware partners for both training and inference workloads. Jalapeño is therefore positioned as an additional part of OpenAI’s computing infrastructure rather than an immediate replacement for Nvidia hardware.
OpenAI also describes Jalapeño as the beginning of a multi-generation hardware platform. The company is continuing production qualification, software development and testing as it prepares to operate the chip at scale.
For now, Jalapeño is an inference-focused accelerator rather than a general-purpose replacement for Nvidia’s full AI hardware ecosystem. But its early benchmark results show why custom silicon could become an increasingly important part of the AI industry’s race for faster and more efficient computing.




