8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

Google Just Dropped Gemini 3.7 Flash and It's Shockingly Impressive

9.4/10
AIAI RevolutionAugust 14, 2026 at 10:09 PM14:33
Audio player
0:00 / 0:00

TL;DR

Google has launched Gemini 3.7 Flash just three weeks after 3.6 Flash, betting on cheaper, stronger coding and agent performance as AI competition shifts from raw capability to cost and speed.

KEY POINTS

Faster release cycle, sharper focus

Gemini 3.7 Flash arrived on August 13, only three weeks after Gemini 3.6 Flash on July 21. Google DeepMind describes it as its strongest workhorse model so far, with improvements aimed squarely at coding, tool use, and agentic workflows rather than general knowledge or math.

Coding benchmarks jumped sharply

On Frontier Code 1.1 Main, 3.7 Flash scored 43.6%, up from 34.4% for 3.6 Flash. On DeepSWE 1.1, which measures longer software engineering tasks, it rose from about 49% to 65.3%, suggesting unusually rapid gains in production-style coding over a short period.

Agent performance improved across the board

The biggest gains came in execution-heavy tests. On Terminal Bench 2.1, 3.7 Flash reached 85.8% versus 78.0% for 3.6; on Terminal Bench 3.0, it climbed from 5.4% to 14.9%; on Automation Bench, from 17.0% to 30.4%; and on OS World 2.0, from 33.8% to 47.9%. The emphasis is on finishing multi-step jobs, not just producing a strong first answer.

Web app generation also advanced

Google says the model is better at building more complete web apps from fewer prompts, reconstructing interfaces from screenshots or images, and keeping a design system consistent through longer builds. On Web Arena, its rating moved from 1,538 Elo to 1,588.

Competitive with top rivals in several tests

On the Artificial Analysis Intelligence Index, Gemini 3.7 Flash scored 56, close to Claude Sonnet 5 at 55 and GPT 5.6 Terra at 57. In coding, 3.7 Flash beat both on Frontier Code 1.1, though GPT 5.6 Terra still led on DeepSWE 1.1 with 69.6% and on Terminal Bench 3.0 with 20.8%.

Pricing is aggressively promotional

Through December 31, 2026, 3.7 Flash is priced at $0.75 per million input tokens and $3.75 per million output tokens, half the original price of 3.6 Flash. From January 1, 2027, prices are set to return to $1.50 input and $7.50 output. The discount reflects a push to win market share in agent workloads, where token and tool-call costs can escalate quickly over many steps.

Already deployed across Google products

The model was put into Gemini Spark on launch day for AI Pro and Ultra subscribers in more than 160 countries and regions. It is also available through the Gemini API, Google AI Studio, Android Studio, and enterprise agent platforms, with use cases centered on Google Workspace actions such as combining files, drafting emails, and updating project documents.

Launch comes amid leadership upheaval

The release follows a major restructuring at Google DeepMind on August 5. Demis Hassabis moved from chief executive to chairman and Alphabet chief scientist, while Koray Kavukcuoglu took over day-to-day operations and now leads Gemini research and development. Sundar Pichai reportedly holds final decision-making authority, underscoring pressure to accelerate flagship model progress.

Flagship delays remain unresolved

Despite the quick Flash updates, there is still no public date for Gemini 3.5 Pro. Google had said in July that 3.5 Pro was in partner testing and that Gemini 4 was in its most ambitious pre-training run yet. The absence of fresh timing has drawn attention as rivals push hard on different fronts.

Rivals are attacking speed and pricing from other angles

OpenAI has previewed Ultrafast, a new API serving tier for GPT 5.6 Soul that reportedly delivers up to 750 output tokens per second, around 14 times normal speed, using Cerebras wafer-scale chips. DeepSeek, meanwhile, released V4 Pro 0813 at $1.32 per million input tokens and $3.96 output, substantially above its V4 Flash pricing, as it seeks to monetize stronger agent capabilities while expanding aggressively.

CONCLUSION

The latest AI model contest is increasingly defined by who can deliver capable agents at the best cost and speed, not just who posts the highest headline benchmark. Gemini 3.7 Flash shows Google pushing hard on that economic and operational battleground while larger questions about its delayed flagship models remain open.

Explain this
Full transcript

More from AI