Google DeepMind shipped Gemini Robotics 2.0, marking a significant upgrade to its physical AI platform that powers general-purpose robots. The release includes three new sub-models designed to give robots whole-body intelligence and improved dexterity for complex humanoid hands.
The centerpiece is Gemini Robotics ER 2, an upgraded embodied reasoning model that processes live video feeds from robot cameras. This vision language model can classify video frame completeness with 60% accuracy and identify key task moments with nearly 90% precision — crucial for knowing when to stop pouring coffee or adjusting grip on a rolling ball.
ER 2 integrates with the Gemini Live API and is publicly available for developers starting today. The model enables real-time failure detection, allowing robots to retry individual steps rather than restarting entire task sequences.
Multi-robot coordination breakthrough
The new system demonstrates multi-robot collaboration through video demos showing Apptronik's Apollo 2 and Franka F3 Duo working together without interference. Google says the robots appear less hesitant than previous versions, though they still lag human speed and grace.
After mapping tasks, ER 2 hands control to Gemini Robotics 2, a vision-language-action model that generates robot movements like other generative AI creates text. A low-latency offline version called Gemini Robotics On-Device 2 can adapt to new robot designs with just 200 movement examples.
Google tested the platform on Boston Dynamics hardware alongside its existing robot partners. The action models show improved accuracy and efficiency compared to the previous 1.6 release from April.
Safety measures expanded
Google introduced ASIMOV-Agentic, a new safety benchmark evaluating whether robots refuse unsafe commands, assess task safety, and request human assistance when uncertain. The company says ER 2 demonstrates robust ability to halt actions when humans approach too closely.
The safety framework combines traditional physical safeguards with AI safety measures across each system layer. Google made the complete ASIMOV-Agentic benchmark available on Hugging Face for broader research use.
The models remain limited to select testers, with broader availability planned as Google continues developing what DeepMind researchers call "physical AGI" — robots capable of any human task through natural language commands.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.