Daily Podcast full article
OpenAI benchmarks Jalapeño ASIC: first numbers put inference efficiency in focus
OpenAI used Hot Chips 2026 to turn Jalapeño from a custom-chip promise into a measured inference platform, claiming higher work per watt and lower latency than leading commercial systems while acknowledging that production scale, independent replication and Nvidia’s next roadmap remain open tests.

A benchmark disclosure aimed beyond the lab
OpenAI’s Jalapeño ASIC now has public performance numbers, and the timing is as important as the silicon. At Hot Chips 2026, the company disclosed first measured results for its clean-sheet inference accelerator, positioning the chip as working first-party hardware rather than a speculative roadmap item . The announcement does not mean Jalapeño is about to appear on a price sheet; it means OpenAI wants cloud partners, investors, chip suppliers and rivals to evaluate whether its own silicon can change the economics of serving AI.
The headline claim is straightforward: across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than comparison systems on SemiAnalysis’ InferenceX benchmark . For highly interactive workloads, OpenAI says the advantage rose to 2.1 to 4.1 times higher performance . Those are not abstract FLOP claims; they target the serving metrics that matter when millions of users are waiting for responses.
What OpenAI actually tested
OpenAI framed the results around matched user experience rather than peak chip performance alone. The company said InferenceX measures the full process of serving an AI request, and OpenAI normalized comparisons by each accelerator’s published chip power rating . Jalapeño is rated at 700 watts, while OpenAI said measured sustained power stayed at or below 550 watts on the tested workloads .
The model mix was also deliberate. GPT-OSS 120B tested OpenAI’s own open-weight family, while DeepSeek R1 670B and Moonshot AI’s Kimi K2.5 1T were external models . That matters because OpenAI is trying to present Jalapeño as a general-purpose AI inference accelerator, not a hard-wired engine useful only for ChatGPT. EE Times reported that Richard Ho, OpenAI’s vice president of hardware, described the part as a ground-up accelerator for large-language-model workloads, developed with Broadcom and Celestica and aimed at reducing data movement .
The most detailed public figures show why latency is central to the story. On GPT-OSS 120B, OpenAI reported about 1.9 times higher peak mixed tokens per second per kilowatt than a GB200 comparison and lower end-to-end latency of 1.03 seconds versus 1.80 seconds . On DeepSeek R1 670B, the comparison used GB300, and OpenAI reported about 1.7 times higher peak mixed throughput per kilowatt and 1.65 seconds of end-to-end latency versus 5.99 seconds . On Kimi K2.5 1T, OpenAI reported about 1.5 times higher peak mixed throughput per kilowatt and 1.56 seconds of latency versus 5.31 seconds .
The architectural bet: less movement, more locality
Jalapeño’s message is not simply “more compute.” OpenAI says the architecture was built around the reality that inference has different phases: prefill is compute-intensive, while token-by-token decode is often constrained by memory bandwidth and communication delays . The chip is designed so model state, including the KV cache used during generation, can be explicitly placed and kept local while compute, memory and networking are activated for each phase .
That focus is consistent with outside reporting. Tom’s Hardware noted that each Jalapeño package pairs the compute die with six HBM4 stacks, totaling 216 GiB and 15.4 TB/s of memory bandwidth, and said OpenAI’s presentation emphasized exposing aggregate HBM bandwidth rather than merely adding more of it . Data Center Dynamics reported that OpenAI plans racks with 128 Jalapeño chips and that a full pod would contain 2,048 ASICs; the 128-chip deployment was described as capable of 1.7 exaflops of 4-bit compute with 27.5 TB of HBM4 .
The system-level framing is crucial. If Jalapeño can keep more data local and avoid unnecessary movement of weights and KV cache, the chip can improve latency without sacrificing throughput. That is exactly the trade-off cloud AI operators care about as chatbots evolve into agents that take many sequential steps.
Why the benchmark still needs caveats
The results are notable, but they are not a final verdict on the AI accelerator market. SemiAnalysis said it benchmarked Jalapeño with OpenAI engineers in the lab, but also cautioned that the full InferenceX suite was not run independently and that AgentX, its preferred benchmark for long-context, multi-turn agentic workloads, had not yet been shown . That caveat matters because production agent traffic stresses routing, prefix cache behavior, cache management and offload systems in ways a single-turn 8k/1k workload may not .
There is also the comparison problem. The main public comparisons put Jalapeño against Nvidia GB200 and GB300 systems, not Nvidia’s Vera Rubin platform . TechCrunch reported that Ho estimated Jalapeño would deploy at the end of 2026 in very small volumes, with more meaningful deployment in 2027, which means the competitive target may shift by the time OpenAI has scale . Tom’s Hardware similarly noted that Vera Rubin was not in the comparison and that Jalapeño is not designed for training, where Nvidia remains central to OpenAI’s infrastructure .
In other words, Jalapeño’s first numbers are a strong opening argument, not a market share result. They show that OpenAI can produce credible first-generation inference silicon, but they do not yet show cost at scale, yield, full production reliability, software maturity across frontier workloads or performance against every next-generation competitor.
Strategic pressure on Nvidia and the cloud stack
Even with caveats, the disclosure changes the conversation. Axios reported that OpenAI does not plan to offer Jalapeño to others because it expects to need the capacity internally . That makes Jalapeño less like a merchant chip and more like a vertical-integration tool: OpenAI can use it to reduce serving cost, improve latency and give itself a bargaining alternative when buying external accelerators.
That alternative does not eliminate Nvidia. Axios quoted OpenAI’s position that Jalapeño is part of a broader compute mix that also includes Nvidia, AMD and other systems, and TechCrunch noted that the chip is aimed at inference rather than training . But it does signal that model companies no longer want to be purely downstream customers of semiconductor vendors. They want to shape memory, networking, compilers, kernels and racks around their own traffic patterns.
OpenAI’s own companion essay made that business logic explicit, arguing that better hardware designed for its workloads improves speed and efficiency, while a broader portfolio of partners preserves choice and pricing discipline . That is the strategic subtext behind Jalapeño: the benchmark is not just about a chip, but about who captures margin and control in the AI stack.
The next test is deployment
OpenAI says it plans to begin deploying Jalapeño within its compute infrastructure by the end of 2026, with Gen 2 already deep in development and Gen 3 taking shape . Data Center Dynamics reported that limited quantities are expected later this year and broader deployment next year . If those timelines hold, the industry will soon learn whether the benchmark lead survives real traffic, real failure rates and real supply constraints.
For now, Jalapeño gives OpenAI a credible first proof point. It shows that a model company can design an inference ASIC around its own workloads, use public benchmarks to challenge incumbent accelerators and make efficiency a strategic lever. The harder question is whether this first disclosure becomes a fleet-scale advantage before Nvidia, AMD, Google, Amazon and Microsoft’s custom-silicon teams move the frontier again.
Sources from the last 72 hours
- [1]OpenAI Inc. (via Public) / Jalapeño’s first results show industry-leading speed and efficiency in AI inferenceAug 25, 2026, 8:28 AM UTC
- [2]OpenAI Jalapeño: Better Than Nvidia BlackwellAug 25, 2026, 10:05 PM UTC
- [3]OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks showAug 25, 2026, 2:22 PM UTC
- [4]OpenAI details Jalapeño AI chip, with 700W TDPAug 25, 2026, 12:00 PM UTC
- [5]OpenAI’s 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU — claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with BroadcomAug 25, 2026, 6:05 PM UTC
- [6]OpenAI says its Jalapeño chip offers spicy performanceAug 25, 2026, 2:00 PM UTC
- [7]First Benchmarks Revealed for Jalapeño, OpenAI’s Clean-Sheet General Purpose AI Accelerator ASICAug 27, 2026, 12:00 PM UTC
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.

Comments
Be the first to comment.