8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

What's new in the Gemma open model family

10/10
GoogleGoogle for DevelopersMay 22, 2026 at 05:59 PM47:46
Audio player
0:00 / 0:00

TL;DR

Google’s Gemma 4 open-weight AI models expand local, efficient, and customizable AI from edge devices to cloud systems, with major gains in performance, multimodality, and deployment flexibility.

KEY POINTS

Expanded model range and efficiency focus

Gemma 4 introduces four model sizes from 2B to 31B parameters, designed to run across devices from IoT hardware to high-end GPUs. The lineup includes a mixture-of-experts 26B model optimized for runtime efficiency and a 31B model for maximum quality and fine-tuning. The models aim to deliver high “intelligence per watt,” matching or exceeding larger predecessors despite smaller sizes.

Major performance leap over previous generation

A key milestone is that the 2B model now rivals or surpasses last year’s 27B model in benchmarks, signaling rapid efficiency gains. Across tasks such as reasoning, coding, and creative writing, improvements are reported across the entire model family, with competitive results against significantly larger systems.

Longer context and agent-ready capabilities

Context windows have expanded from 32,000 to up to 256,000 tokens, enabling more complex tasks and longer interactions. All models now include reasoning (“thinking”), function calling, multi-step planning, and tool use, positioning them for autonomous agent workflows and enterprise automation.

Shift to Apache 2.0 licensing

The move to an Apache 2.0 license replaces earlier custom terms, allowing broader commercial use and easier integration into production systems. This change has been widely welcomed as it simplifies enterprise adoption and ecosystem growth.

Strong multimodal capabilities

Gemma 4 enhances vision, audio, and multilingual processing. Models support variable image formats, object detection with bounding boxes, document and chart understanding, and built-in multimodal translation. Audio capabilities include speech recognition, transcription, and translation, even on smaller models.

Global language performance

Benchmarks show strong multilingual results, with the 31B model ranking among the top across European languages and performing competitively in Japanese, Korean, and Southeast Asian languages. The models approach the performance of leading proprietary systems in several language benchmarks.

High reasoning and tool-use benchmarks

In complex reasoning tests such as BFCL, the 31B model competes with models exceeding one trillion parameters, demonstrating strong planning and execution abilities, including error correction and iterative problem solving.

Local-first deployment and edge AI

Gemma 4 is designed to run locally across devices, including smartphones, browsers, and IoT hardware. Demonstrations include real-time assistants, offline multimodal processing, and robotics applications running on devices like Raspberry Pi and Jetson Nano.

Hybrid and scalable cloud deployment

On Google Cloud, deployment options range from serverless APIs to fully managed endpoints and Kubernetes-based infrastructure. Developers can choose between per-token pricing or dedicated endpoints, with support for fine-tuning and reinforcement learning.

Speed improvements with speculative decoding

The introduction of MTP (speculative decoding) delivers up to 3× faster inference speeds, improving responsiveness for real-time applications and multi-agent systems.

Growing ecosystem and adoption

The Gemma ecosystem has surpassed 500 million downloads and includes over 100,000 fine-tuned variants. These range from domain-specific models like MedGemma for healthcare to language-specialized versions improving accessibility in regions such as Africa and Eastern Europe.

Real-world applications and accessibility

Demonstrations highlight practical uses such as AI agents optimizing transportation systems, offline mobile assistants, and tools for visually impaired users that provide real-time navigation guidance. These examples emphasize on-device AI’s role in privacy, reliability, and accessibility.

CONCLUSION

Gemma 4 positions open-weight AI as a scalable, efficient alternative to large proprietary systems, enabling advanced multimodal and agent-driven applications across local devices and cloud environments.

Explain this
Full transcript

More from Google