8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

OpenAI’s New AI Chip Just Got Real (Beats NVIDIA)

9.3/10
AIAI RevolutionAugust 28, 2026 at 01:53 AM15:56
Audio player
0:00 / 0:00

TL;DR

OpenAI says its new Jalapeno inference chip outperforms Nvidia GB300-class accelerators on public efficiency and latency tests, while fresh Anthropic rollout rumors and Alibaba’s cheaper Qwen release show how fast the AI inference market is shifting.

KEY POINTS

OpenAI posts eye-catching efficiency claims

OpenAI published benchmark results for Jalapeno, a custom ASIC built with Broadcom for inference rather than training. In a public InferenceX benchmark, the company said the chip delivered between 1.5x and 1.9x more AI work per watt than top recorded Nvidia GB200/GB300 results at peak throughput, while also cutting end-to-end latency.

The 104x figure comes with a major caveat

The headline-grabbing result was a test that matched Jalapeno to the rival system’s fastest decoding speed, then compared throughput per kilowatt. Under that framing, Jalapeno reached 104.3x the throughput per kilowatt on DeepSeek R1, plus 53.7x on GPT OSS 120B and 56.1x on Kimi K2.5. That does not mean the chip is 100 times faster overall; it reflects how sharply the rival system’s efficiency drops when pushed to extreme interactive decoding speeds.

Power normalization is central to the comparison

OpenAI normalized results by power, not by chip count. Jalapeno is rated at 700 watts, versus 1,200 watts for GB200 and 1,400 watts for GB300, and sustained measured power was reportedly at or below 550 watts on the tested workloads. The company argues that performance per kilowatt is the more relevant metric for data center deployment, but the choice also heavily shapes the final ratios.

Latency improved alongside throughput

On highly interactive workloads, Jalapeno delivered between 1.7x and 3.6x higher performance at about six times lower end-to-end latency. Full request latency on GPT OSS was 1.03 seconds versus 1.80, on DeepSeek R1 1.65 versus 5.99, and on Kimi K2.5 1.56 versus 5.31. Raw per-user decoding speed was also much higher, reaching 1,459 tokens per second on GPT OSS.

The benchmark used public-weight models

The tests were run on GPT OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T rather than proprietary frontier systems. That choice makes the setup more reproducible and avoids asking the market to trust hidden model behavior. OpenAI said internal tests on its own frontier models showed an even wider advantage, but did not publish those numbers.

The design targets inference bottlenecks directly

The chip is built around the split nature of large-model inference: prefill is compute-bound, while decode is limited by memory bandwidth and communication overhead. OpenAI said Jalapeno keeps model state and KV cache local where possible and treats networking as part of the architecture, allowing it to shift more smoothly between prefill-heavy and decode-heavy workloads.

AI helped design and optimize the chip

OpenAI said models were used to explore implementations, shorten verification loops, and optimize arithmetic circuits, helping move from initial design to tape-out in nine months. Using Codex with GPT Astra, engineers reportedly adapted three open-weight models to the chip in two months, and AI-generated kernels for selected blocks in GPT OSS ran 1.5x to 1.8x faster than human-written versions.

Deployment remains gradual

Small-volume deployment is planned by the end of this year, with a ramp through 2027. OpenAI said the chip is not a replacement for Nvidia, and that its broader compute strategy will continue to rely on external partners for both training and inference.

Anthropic faces rollout rumors and product criticism

Two temporary model names, Melon and Marshmallow, appeared briefly and then vanished, fueling speculation that Fable 5.1 and possibly Sonnet 5.1 are being silently rolled out. Testers reported cleaner front-end code, better long-chain reasoning, stronger tool use, and improved legal drafting, but also stricter copyright refusals. At the same time, Opus 5 has drawn complaints for being stubborn, lazy, and prone to apology loops after errors; Anthropic has acknowledged consistency problems and said fixing that behavior is a priority.

Alibaba cuts training cost with Qwen 3.8 Flash

Alibaba released Qwen 3.8 Flash, a multimodal model positioned as stronger on coding and office tasks than Qwen 3.7 Plus at roughly one-ninth the training cost. It offers a default context window of 262,144 tokens, expandable to 1 million, with pricing of 1 yuan per million input tokens and 3 yuan per million output tokens. The company also open-sourced Qwen 3.8 Flash Next as a preview of architecture planned for the Qwen 4 family.

Claude memory now carries across products

Anthropic has merged memory across chat and Claude Code, reducing the need to re-brief the system when moving projects between tools. Users can view, edit, and delete stored memory, while sensitive categories such as health, religion, politics, race, and gender identity remain excluded by default unless users opt in.

CONCLUSION

The week’s developments point to a market increasingly defined by inference efficiency, lower operating cost, and tighter product integration. The biggest takeaway is not a single benchmark ratio, but the accelerating competition to control how AI is served at scale.

Explain this
Full transcript

More from AI