ENFR
8news

Tech • IA • Crypto

TodayShortsTop StoriesTopicsAll videosYT channelsCryptoArchivesFavorites

Introducing Gemini Robotics 2

9.1/10
GoogleGoogle for DevelopersJuly 31, 2026 at 07:00 PM39:06
Audio player
0:00 / 0:00

TL;DR

Google DeepMind has unveiled Gemini Robotics 2, a new generation of models aimed at bringing general-purpose intelligence, dexterity, and collaboration to robots.

KEY POINTS

A push toward general-purpose robotics

The new Gemini Robotics 2 suite is designed as an “intelligence layer” that can power many types of robots across tasks and environments. The goal is to move beyond narrow automation toward systems that understand surroundings, reason about objectives, and execute actions with human-like adaptability. This marks a shift from task-specific programming to broadly capable robotic intelligence.

Three core capabilities introduced

The release focuses on three advances: whole-body reasoning, improved dexterity, and multi-robot collaboration. Robots can now interpret their full physical state in space, enabling complex actions like cleaning cluttered environments. They are also better at manipulating objects and can coordinate with other robots to divide or parallelize tasks.

Dexterity remains the hardest challenge

Despite progress in locomotion, dexterous manipulation is still a major bottleneck. Tasks like folding laundry or handling flexible objects require managing many contact points and coordinating over 20+ degrees of freedom in robotic hands. These challenges make everyday actions far more complex than movement alone.

Data scarcity limits progress

A key constraint is the lack of large-scale datasets capturing physical interactions. Unlike text or images, there is no “internet of motion” documenting how forces, touch, and movement interact in real-world tasks. High-quality data from teleoperation is effective but expensive, while video data lacks crucial action labels such as force and intent.

Vision-Language-Action models drive capability

The system builds on Vision Language Action (VLA) models, which combine visual input, natural language understanding, and physical control. This allows robots to interpret commands like “clean the garage” or “pick up pink objects” without precise coordinates, a major leap from earlier robotics systems.

Cross-robot generalization emerges

Gemini Robotics 2 is trained across multiple robot types, enabling early signs of cross-embodiment transfer. Skills learned with one form factor, such as grippers, can partially transfer to more complex hands. This flexibility is crucial as hardware designs continue to evolve rapidly.

Industrial deployment likely comes first

Early real-world use is expected in industrial and semi-structured environments, where conditions are more controlled and safety can be managed. Home deployment remains more challenging due to unpredictability, safety risks, and the need for higher reliability in human-centric spaces.

Safety and human interaction prioritized

The models include improvements in human detection, proximity awareness, and task uncertainty. They can identify ambiguous instructions and request clarification, reducing risk. A new benchmark, Asimov, evaluates decision-making in real-world physical scenarios involving safety and common sense.

Integration with broader AI systems

The models leverage advances from large-scale AI systems like Gemini, inheriting strengths in language and vision. However, physical reasoning—such as force and motion—still requires dedicated training, as it does not naturally emerge from digital data alone.

Timeline for widespread adoption

Experts estimate 5 to 10 years before general-purpose robots become common in daily life, a timeline that has shortened significantly due to rapid progress. However, full integration into society will also depend on solving safety, hardware sensing, and human interaction challenges.

CONCLUSION

Gemini Robotics 2 represents a significant step toward versatile, intelligent robots, but key hurdles in dexterity, data, and real-world deployment must still be overcome before widespread adoption becomes reality.

Full transcript

More from Google