OpenAI's first custom inference chip, known as Jalapeño, outperformed Nvidia Corp.'s GB300 processor in testing, chip chief Richard Ho said in a Bloomberg interview.
Ho said the chip outperfomed in tests measuring AI performance per watt and response speed.
“In the lab, Jalapeño is showing performance both in the high-throughput domain, meaning it will be able to serve a lot of customers more cheaply, as well as the low-latency domain, meaning that for the customers that care about it, the response time will be really, really fast,” Bloomberg quoted Ho in a report on Tuesday.
Jalapeño was developed in partnership with Broadcom Inc., which makes custom chips for a range of clients. The two companies announced their collaboration last year and said in June the processor was developed in record time.
The chip is built for AI inference, which involves trained AI models responding to prompts and handling tasks. It is not designed for training AI models, where Nvidia’s technology excels, Ho said.
Jalapeño was also not tested against Nvidia’s newest Vera Rubin chips, which recently began shipping.
Tests
OpenAI said in a blog post that it tested Jalapeño on InferenceX, a public benchmark from SemiAnalysis that measures how quickly and efficiently a system can serve an AI request. The tests covered both high-throughput and low-latency workloads.
On DeepSeek R1 670B, Jalapeño delivered 1.7 times the peak mixed-token throughput per kilowatt of the Nvidia GB300 system. Its end-to-end latency was also 3.6 times lower, at 1.65 seconds versus 5.99 seconds for the GB300 system.
Kimi K2.5 1T, the largest public model in the test, produced a 1.5-times gain in peak performance per watt for Jalapeño, OpenAI said. End-to-end latency was 3.4 times lower.
OpenAI also compared Jalapeño with Nvidia’s GB200 in GPT-OSS 120B. The chip delivered 1.9 times higher peak mixed-token throughput per kilowatt, at 85,448 versus 44,960.
