Google DeepMind is launching Gemini Robotics ER 2, its most capable embodied reasoning model for robotics, on July 30, 2026. The company frames the model as a high-level brain that lets robots chat with people, understand the physical world, and plan multi-step tasks while handing motor execution to lower-level vision-language-action models.
The design lets a robot reason about the next step while it keeps moving, so Gemini Robotics ER 2 can command action models and robotics APIs through multi-step work without stop-and-think pauses. Google DeepMind demonstrated that fluid orchestration on Boston Dynamics Spot by driving navigation and manipulator APIs so the robot fetches objects on natural-language request.
Developers build the same agentic pattern by declaring low-level control interfaces such as VLA models or navigation APIs as tools and streaming video, audio, or text straight into the model.
Continuous video understanding gives the system real-time progress tracking. By quantifying how far a task has advanced, Gemini Robotics ER 2 supplies situational awareness so robots can adjust mid-run or retry a failed step without restarting the full workflow.
On moment-finding benchmarks the model reaches 91.3 percent accuracy with a 0.96-second mean absolute distance, letting it time precise switches such as when to stop pouring. Progress classification across five completion bands hits 57.4 percent accuracy in Google DeepMind’s evaluations.
Gemini Robotics ER 2 also coordinates multiple machines. Diverse robots share a semantic understanding so they can hand off work and finish jobs no single platform could handle alone; Google DeepMind shows Apptronik’s Apollo 2 humanoid collaborating with a Franka F3 Duo arm.
Spatial upgrades now run success and failure detection on raw video rather than still frames, catching spills, slips, or misalignments as they happen, and expand instrument reading across ten display types.
Safety gains appear on instruction-following and human-proximity tests. The model halts a humanoid when a person enters the workspace and resumes only after the area clears.
Google DeepMind states most physical-world tasks remain complex and require multiple steps, so the model’s value sits in orchestration rather than single-shot motion. Gemini Robotics ER 2 is publicly available to developers through the Gemini API and Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform, and configuration examples plus GitHub code accompany the release.













