Google has recently introduced Gemini Robotics 2, a series of AI models that are propelling us forward in the realm of intelligent robots. Gemini Robotics ER2 serves as the central intelligence for one or more robots, empowering them to plan and collaborate effectively. This innovation marks a significant milestone as it unlocks intelligent whole-body control, advanced dexterity, and multi-robot cooperation, shaping the next generation of adaptable robots.
Despite the rapid advancements in foundation AI models for text and code in recent years, the journey of applying AI to physical hardware has been more extensive for Google DeepMind. Over nearly a decade of dedicated work, Google has progressed from initial reinforcement learning experiments in simulations to the development of its Robotic Transformer (RT-1 and RT-2) research. In March 2025, Google launched the first Gemini Robotics family based on Gemini 2.0, transitioning from single-purpose policies to multimodal foundation models capable of translating visual inputs directly into robotic actions.
By collaborating with third-party robot manufacturers like Boston Dynamics, Google incorporated Gemini’s multimodal intelligence into robots like Atlas, enhancing their ability to carry out complex and adaptive tasks in various work settings. The field of robotics has witnessed the successful adoption of reinforcement learning for mastering physical dynamics, enabling robots to excel in tasks such as sprinting, regaining balance, and intricate manipulations.
Gemini Robotics 2 signifies a departure from solely relying on end-to-end reinforcement learning for physical AI. Google DeepMind has embraced a two-tier architecture powered by vision-language foundation models (VLMs): The High-Level Cognitive Layer (VLM) processes visual and audio input for reasoning and task management, while the Low-Level Execution Layer (VLA / Motor Policies) coordinates real-time movements and manipulations based on higher-level directives from the VLM.
The release of Gemini Robotics 2 segments physical AI into three components – Gemini Robotics ER 2 for cognitive processing, Gemini Robotics 2 for cloud-based execution, and Gemini Robotics On-Device 2 for local processing, enhancing the overall intelligence and adaptability of robots in various environments. This advancement enables robots like Atlas to be ready for production this year, integrating generative AI capabilities for improved performance and versatility.
Moreover, to keep up with the latest developments in technology, sign up for the weekly newsletter, subscribe to the RSS feed, or follow on social media platforms. Stay informed about new releases, such as the free Python course for beginners offered by Scrimba, and explore the latest updates on Apache Fory’s capabilities in JSON for Java.
