OpenAI Shares Fast Inference Benchmark Results for Its Jalapeño Chip

Detailed at the Hot Chips conference, the Jalapeño chip demonstrated high performance and efficiency compared to existing inference processors.

◉ 0 views
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show | TechCrunch

OpenAI shared detailed benchmark results for its new Jalapeño chip with attendees at the Hot Chips conference, announcing that the chip has been developed for scalable, fast inference.

Test Results Shared at the Conference

OpenAI publicly announced the first benchmark results for its new Jalapeño hardware system at the Hot Chips conference held on Tuesday.

In InferenceX tests conducted by SemiAnalysis, the chip achieved higher efficiency figures compared to the most advanced existing inference processors.

Performance and Power Efficiency Advantages

Richard Ho, head of hardware at OpenAI, stated that the conducted tests pointed to a massive performance advantage compared to existing technology.

It was explained that Jalapeño can execute more AI workloads per unit of power and return responses to users with much lower latency.

Competitive Landscape and Shipment Schedule

While current comparisons are made against the Nvidia Blackwell system, it remains to be seen how the market will shape up by the time Jalapeño is fully deployed in the field.

Richard Ho stated that shipments of the chip in very small volumes will begin at the end of 2026, while a more comprehensive rollout is planned for 2027.

Broadcom Partnership and Full-Stack Approach

Announced last October, Jalapeño was developed in partnership between OpenAI and Broadcom, with AI models also playing an active role in the process.

The company aims to turn the chip into a multi-generational platform, ensuring that products, models, and memory architecture evolve in a harmonized manner.

Custom Design to Reduce Latencies

Thanks to this integrated approach, the goal is to prevent the stages that frequently create bottlenecks during the inference process.

Designed specifically to minimize latencies especially in the prefill and communication phases, the system keeps data movement to a minimum.

Share