Robotics
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Image: Primary Google announced today the launch of Gemini Robotics ER 2, its most capable embodied reasoning model for robotics. The model functions as a high-level brain that allows robots to chat with humans, understand the physical world, plan multi-step tasks, and hand off motor execution to lower-level vision-language-action models.
Gemini Robotics ER 2 improves on its predecessor by watching continuous video feeds so robots can track progress, adapt to errors, and know when to proceed to the next step. The system introduces multi-robot collaboration, enabling diverse machines to work together in shared spaces and complete complex workflows a single robot could not do alone. It integrates with the Gemini Live API using a bidirectional streaming endpoint optimized for latency-sensitive tasks, eliminating stop-and-think pauses.
The model achieves 57.4% accuracy on progress classification tasks and 91.3% accuracy on moment-finding tasks with a 0.96-second mean absolute distance. It delivers this precision at a fraction of the compute cost and four times the execution speed of larger model categories. Gemini Robotics ER 2 is publicly available to developers via the Gemini API, Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from Google DeepMind and reviewed by the T&B editorial agent team.
