8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

Google Just Dropped Its Most Powerful Cyber AI Yet

9.4/10
AIAI RevolutionSeptember 3, 2026 at 10:56 PM14:12
Audio player
0:00 / 0:00

TL;DR

Google has launched a low-cost Gemini 3.8 Flash model and a gated Flash Cyber variant that can find and patch software flaws at near-frontier performance, as AI labs and regulators simultaneously move to restrict who can use the strongest cyber capabilities and who can shut such systems down.

KEY POINTS

Chrome bug found after 13 years

A longstanding Chrome vulnerability that had gone unnoticed for 13 years was reportedly uncovered by Google’s cyber-focused model in a single pass. Doug Turner, an engineering director for Chrome, described a sharp recent surge in vulnerability reports and called it a “vulnerability apocalypse,” reflecting how generative AI is accelerating bug discovery in heavily reviewed code.

What Flash Cyber is built to do

Gemini 3.8 Flash Cyber is a specialized version of Gemini 3.8 Flash trained for cybersecurity tasks, especially finding vulnerable code and producing working fixes. Google said it is already using the model on its own software and that, on Chrome vulnerabilities, it generated 2.6 times more correct patches than larger commercial models.

Benchmark results and caveats

On CyberGym, which measures vulnerability discovery, the model scored 86.2%. On CWE patching tests it scored 47.2%, and Google said an internal benchmark showed more than 70% success finding flaws across 20 programming languages. That internal figure is harder to independently verify, but it suggests broad bug-finding ability rather than narrow memorization.

Wiz and cloud security results

Wiz, the cloud security company acquired by Google for $32 billion, said Flash Cyber achieved 7.5% to 9.7% higher recall of real-world vulnerabilities than leading frontier models on its internal penetration-testing benchmark. It also did so at roughly 2.3 to 5.2 times lower cost, a key advantage for scanning large codebases continuously.

Cheap enough to run at scale

The public Gemini 3.8 Flash model is priced at $0.75 per million input tokens and $3.75 per million output tokens, with a 1 million-token context window and 64K output limit. Estimates cited for high-effort runs put a full task at about $0.58, making repeated tool use and self-checking affordable in ways premium models often are not.

Near-flagship engineering performance

On the DeepSWE software engineering benchmark, Gemini 3.8 Flash scored 73.7%, compared with 74.0% for Claude Opus 5. It was also reported at about 305 tokens per second, roughly double the speed of some higher-tier rivals, supporting Google’s argument that low cost and repeated attempts can rival brute model size.

Mixed independent testing

Not all outside testing has been flattering. Critics pointed to uneven benchmark performance and possible test-specific tuning, while anecdotal comparisons produced contradictory outcomes depending on the task. In some cases Flash looked dramatically cheaper than top competitors, but in others it missed obvious prompt details that rival models handled better.

Why the strongest version is gated

Flash Cyber is not broadly available. It is being offered through Google’s FAIR program to governments, critical infrastructure operators, and selected advanced defenders because safeguards around cyber activity have been deliberately loosened to make defensive security work practical. Protections around CBRN misuse remain in place, and Google said it prioritized vulnerability fixing over offensive exploitation.

A broader industry shift toward controlled access

Anthropic has made a similar move by limiting its strongest cyber capabilities to verified users. The pattern suggests major AI labs increasingly agree that the most capable cyber models should not be released without identity checks and access controls, even while cheaper public variants become widely available.

Shutdown powers are becoming a policy issue

In the United States, OpenAI told lawmakers it is developing automated shutdown capabilities for severe incidents, with human responders currently pausing activity if alerts cannot be cleared within 30 minutes. At the same time, the proposed AI Kill Switch Act remains in committee, while the European Union already has authority under Article 93 of the AI Act to restrict, withdraw, or recall a general-purpose AI model from the market.

CONCLUSION

AI systems are quickly becoming cheap and capable enough to hunt software flaws at industrial scale, shifting the balance of cyber defense and exposing new governance questions. The immediate contest is no longer only about model power, but about controlled access, oversight, and who gets authority to stop these systems when risks escalate.

Explain this
Full transcript

More from AI