Google DeepMind's Gemini Robotics 2 controls a whole humanoid body under one policy
Google DeepMind released Gemini Robotics 2, its first model to control an entire humanoid body under a single learned policy.
Google DeepMind released Gemini Robotics 2 on July 30, the first of its models to control an entire humanoid body – legs, torso, arms and multi-finger hands – under a single learned policy.
Gemini Robotics 2 is a vision-language-action model, meaning it converts camera and text input directly into motor commands. The single-policy approach lets a robot walk, crouch, stretch and manipulate objects in one continuous motion, rather than handing off between the separate locomotion and manipulation controllers that earlier systems, including Gemini Robotics 1.5, relied on.
DeepMind shipped two companion models alongside it: Gemini Robotics ER 2, an embodied-reasoning model that plans multi-step tasks and coordinates multiple robots, and Gemini Robotics On-Device 2, which runs locally on robot hardware and adapts to a new robot body with a few hours of data.
Demonstrated tasks include tying knots and sealing ziplock bags with five-fingered hands, plus multi-robot handoffs to finish jobs faster, according to DeepMind.
The company also introduced ASIMOV-Agentic, a safety benchmark for agentic robot systems, and said the ER 2 model improves human-detection and proximity behavior.
The demos come with a caveat DeepMind supplies itself. Its own figures show multi-finger dexterity still ranges widely, from 32% to 92% depending on the task, a reminder that whole-body control does not yet mean reliable hands. The results are the company’s own; there is no independent benchmark of the systems in real-world use.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
