
Tech • AI • Robotics
Google DeepMind says there is no definitive benchmark for AGI, arguing that progress will be measured through increasingly capable, trusted systems built with user feedback, while the company pushes an ambitious Gemini 4 training run and a stronger focus on agentic AI.
Koray Kavukcuoglu, who leads Google DeepMind, said there is no accepted threshold or benchmark that can determine when artificial general intelligence has been reached. He argued that AI progress has historically advanced through shifting goals rather than a single exam, and that what matters most is building systems people can trust and work with across a widening range of tasks.
The path to more general AI is being treated as a process of co-development with users rather than a lab-only milestone. Daily consumer use cases such as email assistance sit alongside scientific research applications, and those interactions are being used to identify which problems are worth solving next and which capabilities matter most in practice.
Recent work on the Gemini family has focused less on raw text generation and more on turning models into agents that can operate with tools, functions and software workflows. Kavukcuoglu said the team learned substantially from work between Gemini 3.5, 3.6 and 3.7, especially in coding and software engineering, where models must collaborate, take actions and manage more complex workflows.
Gemini 4 was described as the most ambitious pre-training effort the group has attempted so far. Kavukcuoglu said the run is progressing well, while stressing a cautious approach until the model proves itself first internally and then with external users.
Kavukcuoglu rejected any suggestion that the group is comfortable operating below the leading edge. He said there is “nothing other than being at the frontier” that matters strategically, and tied that ambition to Google’s computing resources, full-stack optimization and long-term investment culture.
The case for confidence rests partly on Google’s history of investing early in major technical platforms, including AI chips, quantum computing and Waymo. Kavukcuoglu argued that this willingness to fund difficult, long-horizon work gives Google DeepMind unusual capacity to pursue frontier systems at scale.
Even as AI products move into wide public use, Kavukcuoglu said success still depends on maintaining a broad exploration funnel for new ideas. Bringing frontier research and product teams closer together is intended to improve both execution and the flow of innovation into Gemini, without sacrificing the experimentation that may unlock the next leap.
Reflecting on the organization’s history, Kavukcuoglu pointed to early Atari research as a formative moment. The effort that led to DQN showed how deep learning and reinforcement learning could combine to produce agents that learned on their own, a trajectory that later expanded into milestones such as Go, Chess, StarCraft and AlphaFold.
He said the foundations of today’s systems still rely on deep learning, pre-training and reinforcement learning methods that have been used for years. What changed is the environment: language and real-world tasks are far more ambiguous, open-ended and multi-domain than games, demanding models with broader reasoning and better anticipation of user intent.
Within the current Gemini lineup, Pro, Flash and Flash-Lite all remain active tracks, but the faster recent cadence has come from Flash models. Kavukcuoglu said those systems have improved rapidly from 3.5 to 3.7, making them a practical path toward stronger frontier competitiveness even as work continues on 3.5 Pro and Gemini 4.
Google DeepMind is presenting AGI less as a finish line than as a compounding process of capability, trust and real-world use. Its near-term strategy centers on pushing Gemini toward more intelligent, agentic behavior while trying to reclaim and hold the frontier.
Explain this