8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

OpenAI Jalapeno benchmarks, Nvidia-Hugging Face $13B deal rumor

AISaturday, August 29, 2026· 10 videos

Briefing

Audio player
0:00 / 0:00

OpenAI touts Jalapeno efficiency lead

OpenAI published benchmark claims for Jalapeno, a custom inference ASIC built with Broadcom, saying it outperformed recorded Nvidia GB200/GB300 results on the public InferenceX test. The company said the chip delivered roughly 1.5x to 1.9x more AI work per watt at peak throughput while also reducing end-to-end latency. The comparison was normalized by power, with Jalapeno rated at 700 watts versus about 1,200 watts for GB200 and 1,400 watts for GB300. The release sharpens pressure on the inference market, where efficiency and interactive latency are becoming as commercially important as raw training scale.

Jalapeno's 104x claim needs context

The most dramatic figure in OpenAI's data was 104.3x throughput per kilowatt on DeepSeek R1 when Jalapeno was matched to a rival system's fastest decoding speed. Similar power-efficiency spreads were cited at 53.7x for GPT OSS 120B and 56.1x for Kimi K2.5. Those numbers do not mean the chip is more than one hundred times faster overall; they reflect how badly efficiency can deteriorate when other accelerators are pushed to extreme interactive decoding settings. The caveat matters because benchmark framing, not just silicon design, can radically alter the headline.

Nvidia reportedly targets Hugging Face

Nvidia is reportedly close to acquiring Hugging Face for roughly $12.9 billion to $13 billion, a deal that would extend its reach far beyond chips. Hugging Face has become a key hub for open AI development, with about 3 million models, 1 million datasets and 1 million applications in Spaces. If completed, the takeover would give Nvidia influence over a distribution layer often described as the GitHub of AI. The prospect also raises fresh antitrust questions around concentration across hardware, models and developer platforms.

Anthropic expands Claude Co-work browser

Anthropic added a built-in browser to Claude Co-work in its desktop app, allowing Claude to open pages, click, type and read the web from a side panel. The feature reduces dependence on browser extensions and turns the app into a more self-contained work environment. Claude can also be triggered from the web or a phone to drive that desktop browser remotely, as long as the client remains open and connected. Selective cookie import for signed-in sessions makes authenticated workflows possible, while increasing the importance of permission controls.

Claude memory becomes governance issue

A new Claude memory system is rolling out across the app, letting the assistant generate memories from chats, reference previous conversations and even import memory from other AI services. The feature goes beyond profile facts and can capture working habits, style preferences and behavioral signals that shape future outputs. Users can choose whether sensitive areas such as health, politics, religion or intimate details are stored, but remembered context may still influence professional interactions. That makes memory less a convenience setting than a workplace governance and privacy question.

ChatGPT Work pushes deeper automation

ChatGPT Work is being positioned as a multi-app execution layer rather than a standard chatbot, with connectors for Gmail, Slack, Calendar and Google Drive. It can gather information spread across services, draft outputs and return them for user approval instead of requiring manual copy-and-paste between tools. One showcased use case turns a long internal meeting transcript into a structured Google Doc for teams and a corresponding Google Slides deck for leadership. The resulting documents can be saved as reusable templates, signaling a push from ad hoc prompting toward repeatable office workflows.

OpenAI ties AGI to 2026

OpenAI executives said the company could reach its internal definition of AGI before the end of 2026, putting an unusually near-term date on a contested milestone. The system most closely linked to that timeline is Astra, which leaders say can function as an automated AI research intern by implementing experiments in the company's codebase and returning results. Mark Chen reportedly described the organization as about 80% of the way there, while Sam Altman framed the countdown in public terms. None of the central claims has yet been validated externally, leaving the announcement strategically potent but technically unverified.

Seedance 2.5 lifts AI video

ByteDance's Seedance 2.5 and Higgsfield point to a new phase in generative video, with up to 30-second clips, native 1080p output and synchronized audio generation. Seedance 2.5 doubles the 15-second ceiling of Seedance 2.0 and supports as many as 50 references, including images, video clips and audio files. The practical gain is workflow-related: longer generations reduce the need to stitch together many short shots and hide continuity breaks. Even so, scene consistency remains the central weakness, showing that control is improving faster than narrative coherence.

Videos covered

Previous briefings · AI