Updated
Updated · OpenAI · Aug 25
OpenAI Says Jalapeño Chip Delivers 1.5-1.9x More AI Work per Watt
Updated
Updated · OpenAI · Aug 25

OpenAI Says Jalapeño Chip Delivers 1.5-1.9x More AI Work per Watt

3 articles · Updated · OpenAI · Aug 25

Summary

  • 1.5-1.9x more AI work per watt and 1.7-3.6x lower end-to-end latency were recorded for OpenAI’s Jalapeño across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.
  • InferenceX benchmark tests measured matched user experience across throughput and low-latency settings, with OpenAI saying Jalapeño stayed on the Pareto frontier by improving speed, efficiency and responsiveness at once.
  • 550 watts or less of sustained power was measured on tested workloads, versus Jalapeño’s 700-watt rating, while highly interactive workloads showed 2.1-4.1x higher performance.
  • Nine months from design to tapeout, the chip was developed with AI assistance, and OpenAI said AI-generated code sped selected model blocks by 1.5-1.8x over human-written versions.
  • By year-end, OpenAI plans to start deploying Jalapeño in its own infrastructure as the first step in a multigenerational inference-chip roadmap alongside continued use of Nvidia and other partners’ accelerators.

Insights

With Jalapeño cutting inference costs by 50%, will OpenAI's captive AI chip permanently disrupt Nvidia's dominance in the hardware market?
Since Jalapeño is strictly captive silicon, how will independent developers ever match the lightning-fast inference speeds of OpenAI's closed ecosystem?
If AI models helped design OpenAI's new chip in just nine months, could future silicon completely eliminate human hardware engineers?