Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

OpenAI publishes first Jalapeño inference chip benchmarks

OpenAI reported that Jalapeño, its first custom inference chip, delivered up to 1.9 times more AI work per watt and up to 3.6 times lower end-to-end latency than the GB200 and GB300 comparison systems in company-run tests, and said deployment inside its own compute infrastructure begins by the end of 2026.

D
Sep 2, 2026 · 3 min read

OpenAI published the first measured benchmark results for Jalapeño, its first custom inference chip, on August 25, and said it plans to begin deploying the accelerator inside its own compute infrastructure by the end of 2026.

In its results post, the company reported that Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the commercial comparison systems, measured across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. OpenAI produced all of the figures itself, and said production qualification, software maturation, at-scale operating preparation and validation across more models are still underway.

The reported numbers

Per workload, OpenAI reported:

  • GPT-OSS 120B: about 1.9 times higher peak mixed tokens per second per kilowatt and about 1.7 times lower end-to-end latency than a GB200 comparison system.
  • DeepSeek R1 670B: about 1.7 times higher peak mixed tokens per second per kilowatt and about 3.6 times lower end-to-end latency than a GB300 comparison system.
  • Kimi K2.5 1T: about 1.5 times higher peak mixed tokens per second per kilowatt and about 3.4 times lower end-to-end latency than a GB300 comparison system.

Mixed tokens per second per kilowatt counts how many tokens a system produces for every kilowatt it draws, so it credits hardware that serves more requests at lower power rather than hardware that is only faster. OpenAI normalized the comparison using published package power ratings — 700 watts for Jalapeño, 1,200 watts for the GB200 and 1,400 watts for the GB300 — and said Jalapeño’s measured sustained power stayed at or below 550 watts on the tested workloads.

The runs used the public InferenceX benchmark with nominal 8,000-token input and 1,000-token output single-token-prediction workloads, a single-turn shape that does not cover agentic, long-context or multi-turn serving.

What outside observers saw

SemiAnalysis, which publishes InferenceX, said it witnessed InferenceX runs with OpenAI engineers in OpenAI’s lab, but that OpenAI supplied all of the numbers and that it did not run the full InferenceX suite or see AgentX results.

OpenAI did not publish AgentX, long-context, multi-turn or production-traffic results, fleet-scale reliability, utilization or operating-cost data, or a like-for-like comparison against Nvidia’s Vera Rubin systems.

Deployment plans

“We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure by the end of the year,” the company said. In an interview with Axios, OpenAI said it expects only a limited number of Jalapeño systems in 2026 and greater capacity in 2027. “We’re going to need a lot of compute. And Jalapeño is part of that,” Richard Ho, OpenAI’s vice president of hardware, told Axios.

Jalapeño is built for inference — running already-trained models to answer requests — and not for training them. OpenAI said it will continue to deploy Nvidia and other partners’ accelerators broadly, for both training and inference.

OpenAI and Broadcom unveiled Jalapeño in June as OpenAI’s first inference processor, with initial deployment designed for the end of 2026. OpenAI designed the chip, while Broadcom contributed silicon implementation, networking and connectivity, and Celestica contributed board, rack and system expertise. The two companies announced in October 2025 a collaboration to deploy 10 gigawatts of OpenAI-designed accelerator systems, with deployment targeted to start in the second half of 2026 and finish by the end of 2029.

More news