Google releases the strongest embodied reasoning model Gemini Robotics ER 2: supports continuous video understanding and multi-robot collaboration

Yesterday, Google DeepMind published a blog post, announcing the launch of a new AI model specifically designed for robotics, claiming it enables humanoid robots to coordinate whole-body movements.
The new model, named Gemini Robotics 2, allows robots to perform actions such as walking, squatting, and manipulating objects while executing tasks, combining reasoning capabilities to autonomously complete tasks.
Google stated that unlike previous models that primarily controlled the upper body of robots, Gemini Robotics 2 achieves full-body control of humanoid robots.
In a pre-recorded demonstration, DeepMind's new model controlled Apptronik's humanoid robot Apollo to perform a series of actions, including walking across a room, picking up a watering can and placing it on a shelf, while avoiding obstacles along the way.
At the same time, DeepMind also released two other robot AI models, which can work together or operate independently.
Among them, Gemini Robotics 2 is responsible for converting camera images and natural language instructions into robot motor control commands; while Gemini Robotics ER 2 acts as the robot's "reasoning system," tasked with planning multi-step tasks and coordinating multiple robots to achieve the same goal.
The Gemini Robotics ER 2 model is Google's strongest embodied reasoning model for robotics.
In terms of functionality, compared to the previous version 1.6, the Gemini Robotics ER 2 model, by continuously watching video, now allows robots to track their own progress, make adjustments when problems arise, and accurately determine the timing for the next step.
Additionally, the model can natively call upon Google Search or developer-defined custom functions.
In terms of physical agents, while the robot is currently executing a task, the Gemini Robotics ER 2 model can synchronously reason about subsequent steps to reduce the pauses caused by the traditional "stop-think-act" approach in robots.
This model connects to the bidirectional streaming endpoint of the Gemini Live API, targeting latency-sensitive robot tasks.
Compared to Gemini Robotics ER 1.6, ER 2 can analyze continuous video streams, enabling the robot to track its own task progress, adjust or retry when actions go wrong, and determine when to proceed to the next step.
In the task progress classification test, Gemini Robotics ER 2 achieved an accuracy of 57.4%. This capability helps robots correct failed steps without restarting the entire workflow.
In the moment-finding (key moment localization) test, the model needs to identify the exact video frame where a key event occurs, such as determining when to stop pouring coffee into a cup. ER 2 achieved an accuracy of 91.3% with a mean absolute time error of 0.96 seconds.
Furthermore, Google also demonstrated improvements in robot dexterity. Researchers stated that in tests, the system achieved a 92% success rate for the task of "unscrewing a light bulb."
However, success rates were still relatively low for more complex operations such as tying garbage bags or sealing Ziplock bags.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 4 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 5 hours ago







