OpenAI says its Jalapeño chip beats Nvidia's GB300 on inference

Wait 5 sec.

OpenAI has published its first benchmark results for Jalapeño, its in-house inference chip built with Broadcom.The company claims that the chip does more AI work per watt and returns answers faster than the Nvidia GB200 and GB300 rack systems that it was measured against.How does OpenAI’s new chip perform?OpenAI has published benchmark results for its new in-house AI chip, called Jalapeño. The chip was run through InferenceX, a public benchmark from SemiAnalysis that measures the full job of serving an AI request, and tested it on OpenAI’s own GPT-OSS 120B, DeepSeek’s R1 670B, and Moonshot AI’s Kimi K2.5 1T.OpenAI said its new chip did 1.5 to 1.9 times more work for each unit of electricity than the other systems it was tested against. It also answered requests 1.7 to 3.6 times faster.For the most hands-on tasks that require much back-and-forth, the gap grew to between 2.1 and 4.1 times better.OpenAI’s hardware vice president, Richard Ho, told reporters the chip handles more work at once and also responds quicker, which most chips cannot do at the same time.The comparison points were Nvidia’s GB200 and GB300 superchips, which were the best results InferenceX had on record at the time. On the largest model tested, Kimi K2.5, OpenAI said Jalapeño hit roughly 1.5 times the peak performance per watt and 3.4 times lower latency.There is no rental market, no instance type, and no plan to sell Jalapeno, and when asked whether OpenAI would offer the chip to others, Ho said the company is “struggling to have enough” compute for its own needs.OpenAI has the budget and the floor space, but it is currently limited by data center power rather than budget or floor space, which makes tokens per megawatt the metric that counts.Inference runs continuously across ChatGPT and the API, so shaving watts off each token compounds fast at the scale OpenAI operates. The chip is rated at 700 watts, but held at or below 550 watts on the workloads tested.The chip fits into OpenAI’s October 2025 deal with Broadcom to deploy 10 gigawatts of OpenAI-designed accelerators through 2029. Cryptopolitan reported when the partnership was first sketched out.Will OpenAI switch to using its own chips exclusively?At Hot Chips, Ho said OpenAI will deploy the chip in racks of 128, with a full pod running 2,048 ASICs. A 128-chip deployment delivers 1.7 exaflops of 4-bit compute and 27.5 terabytes of HBM4, with each package carrying 15.4 terabytes per second of memory bandwidth.OpenAI said it used its own models to speed the work on Jalapeno, moving from design to tape-out in just nine months. Ho added that a second-generation version is “deep into development” and a third is underway.However, even with the benchmark claims, OpenAI is not swapping out its GPU fleet. Ho described Jalapeño as one piece of a compute strategy that still leans on “very, very good partners” at Nvidia and Cerebras. The chip will reportedly be deployed in small volumes by the end of 2026 before ramping through 2027.Notably, all the numbers came from OpenAI; SemiAnalysis verified the InferenceX runs in person but did not run the full suite or see results from AgentX. Don’t just read crypto news. Understand it. Subscribe to our newsletter. It's free.