

OpenAI unveiled results for Jalapeño, its custom AI processor developed with Broadcom, on August 25, 2026. The OpenAI Jalapeño chip targets faster AI inference through specialized hardware, memory and networking.
OpenAI plans deployment inside its computing infrastructure by the end of 2026. The company says the design can deliver more AI work per watt while reducing response latency.
The custom AI processor focuses on large language model inference instead of general AI workloads. OpenAI tested Jalapeño across GPT OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The tests used InferenceX, a public benchmark from SemiAnalysis, against commercial AI systems.
OpenAI reported 1.5 to 1.9 times more AI work per watt at peak throughput. The chip also showed 1.7 to 3.6 times lower end to end latency. Interactive workloads recorded 2.1 to 4.1 times higher performance during testing.
Jalapeño tackles inference bottlenecks through tighter coordination between computing, memory and networking. Its architecture keeps model information, including KV cache data, closer to active processing. This approach can reduce data movement and improve response times during model generation.
OpenAI also highlighted Jalapeño's potential for AI agents that perform several tasks sequentially. Lower latency could reduce waiting between planning, tool calls and generated responses. The company expects these gains to support faster products and improved infrastructure economics.
Sam Altman summed up the development on X with a simple statement: “we made a chip and it is fast.”
OpenAI still plans to use Nvidia accelerators alongside Jalapeño for training and inference. The company sees custom silicon as an additional computing layer, rather than an immediate Nvidia replacement.
The bigger shift comes from AI designing AI hardware. OpenAI says its models helped shorten Jalapeño's development cycle to nine months. Gen 2 remains under development, signaling a longer custom-chip strategy for future AI infrastructure.