Basketball Naive Robot Dunks: Google’s Brain Upgrade
- Imagine a world where robots not only understand commands but also adapt to their surroundings, performing tasks with human-like dexterity.
- google DeepMind unveiled its Gemini Robotics and Gemini Robotics-ER models, powered by the gemini 2.0 foundation AI.
- Unlike robots limited to pre-programmed actions, Gemini Robotics can perceive changing surroundings and execute instructions dynamically.
The Dawn of the ‘Physical AI’ Era
Table of Contents
- The Dawn of the ‘Physical AI’ Era
- Physical AI: Your Questions Answered
- What is Physical AI?
- How Does Physical AI Differ from Traditional AI and Generative AI?
- What are the Key Capabilities of Google DeepMind’s Gemini Robotics?
- Can you provide examples of Gemini Robotics in action?
- How does DeepMind measure the performance of AI models for robotics?
- Who are the major players in the Physical AI race?
- What is China’s strategy for Physical AI development?
- What are Vision-Language-Action (VLA) models?
- What is the importance of Nvidia’s Cosmos platform?
- What industries will Physical AI impact?
- Key Companies and Technologies in Physical AI
Imagine a world where robots not only understand commands but also adapt to their surroundings, performing tasks with human-like dexterity. This is the promise of Physical AI, where robots can assess situations and act autonomously.
Google DeepMind’s Gemini Robotics
google DeepMind unveiled its Gemini Robotics and Gemini Robotics-ER models, powered by the gemini 2.0 foundation AI. These models are designed to output physical actions, rapidly adapt to complex real-world environments, and execute tasks effectively. DeepMind stated that these models “understand natural language instructions with greater breadth than previous models and adjust behaviors in response to changes in the surroundings.”
Unlike robots limited to pre-programmed actions, Gemini Robotics can perceive changing surroundings and execute instructions dynamically. In one demonstration, a robot was shown identifying and placing a banana into a bin after a user instructed it to do so. It also dribbled a basketball upon request, showcasing its understanding of physical concepts.
According to Kanishka Rao, a Google DeepMind engineer, ”Although the robot had never seen anything related to basketball before, it understood the shape of a basketball hoop and the concept of ‘dunking’ through the AI model and embodied it in the physical world.”
In another demonstration, a robot followed instructions to fold origami. When asked if it knew that “origami” comes from the Japanese words “ori” (folding) and “kami” (paper), the robot confirmed its awareness. It also performed tasks requiring fine manipulation, such as zipping a backpack and matching the numbers on a die.
deepmind emphasizes that AI models for robotics require “versatility” to adapt to diverse situations, ”situational awareness” to understand changes in the environment, and “proficiency” to perform tasks humans can do. They assert that “Gemini Robotics shows important progress in all three areas.” according to DeepMind’s technical report, Gemini Robotics outperformed most recent Vision-Language-Action (VLA) models in comparative performance benchmark tests, including OpenAI’s Chat GPT 4o and Anthropic’s Claude 3.5 Sonnet.
The Accelerating Race for Physical AI Growth
Global tech companies, including Google, are accelerating the development of Physical AI, aiming to enable robots to assess situations and act accordingly. At CES 2025 in Las Vegas, Nvidia introduced its Cosmos platform for Physical AI development.Jensen Huang, Nvidia’s CEO, predicted a “ChatGPT moment” for robotics. This surge mirrors the global AI boom following the release of ChatGPT.
Microsoft (MS) also entered the fray, releasing its VLA model “Magma” via a research paper last month. Furthermore, Hugging face and Physical Intelligence introduced “Pi0,” an open-source VLA model that transforms natural language commands into autonomous robot actions.
China’s Push for Physical AI
China is also intensifying its efforts in Physical AI development, recognizing its potential to revolutionize industries. The term “embodied intelligence” (å ·èº«æºè½),which defines Physical AI,was included in the government’s work report at this year’s Two Sessions,China’s largest annual political event.
Beijing has outlined concrete policy support, including the “2025-2027 Embodied Intelligence technology Innovation and Industrial Development plan.” The Beijing Humanoid Robot Innovation Center announced the world’s first general-purpose humanoid robot open-source platform on March 12, following the conclusion of the Two Sessions.
According to a professor at Hanyang University, “Chinese robot company UBTECH recently developed a bipedal robot walker, Walker R1.” He added,”It was first introduced to the production line of a Jiko car company as a ‘team unit.'” He further noted that “a supply chain from software development to hardware development to industrialization of humanoid robots is already forming in China, and with government support, technological development and industrialization speed are expected to accelerate further.”
Physical AI: Your Questions Answered
The rise of Physical AI is transforming how robots interact with the world. This Q&A explores the key aspects of this exciting field, from Google DeepMind’s advancements to global competition and applications.
What is Physical AI?
Physical AI bridges the gap between digital intelligence and the real world.Unlike traditional AI, which primarily processes facts, Physical AI empowers machines, robots, and other devices to:
- Perceive their surroundings through sensors.
- Make decisions in real-time based on that sensory input.
- Interact with the habitat through actuators (motors, etc.).
- Adapt their actions based on changing conditions.
In essence, it’s about giving AI a body and the ability to learn and react within a physical space [1], [3].
How Does Physical AI Differ from Traditional AI and Generative AI?
Traditional AI often involves software-based algorithms that analyze data and make predictions. Generative AI, like ChatGPT, focuses on creating new content, such as text or images, based on prompts. Physical AI, however, goes a step further by:
- Embodying intelligence within a physical form.
- Interacting with the physical environment, not just digital data.
- Adjusting its behavior based on real-world feedback, rather than simply answering questions.
What are the Key Capabilities of Google DeepMind’s Gemini Robotics?
Google DeepMind’s Gemini Robotics models, including Gemini robotics-ER, leverage the Gemini 2.0 foundation AI to achieve remarkable capabilities:
- Understanding Natural Language: Interprets instructions with greater breadth.
- Adaptive Behavior: Adjusts actions based on changes in the surrounding environment.
- Physical Task Execution: Capable of performing a variety of tasks, such as identifying objects, manipulating items, and even dribbling a basketball.
- Versatility: Adapts to diverse situations.
- Situational Awareness: Understands and reacts to changes in the environment.
- Proficiency: Performs tasks with human-like dexterity.
Can you provide examples of Gemini Robotics in action?
Gemini Robotics has demonstrated its capabilities in several compelling ways:
- Object Recognition and Placement: Identifying a banana and placing it in a bin upon instruction.
- Understanding Physical Concepts: Dribbling a basketball, showcasing an understanding of spatial relationships and physics.
- Fine Manipulation: Folding origami,zipping a backpack,and matching numbers on a die.
- Knowledge Retention: Confirming its awareness of the Japanese origins of the word “origami”.
How does DeepMind measure the performance of AI models for robotics?
DeepMind emphasizes three key areas when evaluating AI models for robotics:
- Versatility: The ability to adapt to a wide range of situations.
- Situational Awareness: The capacity to understand changes in the environment.
- Proficiency: The skill in performing tasks at a human level.
Gemini Robotics has shown meaningful progress in all three areas, outperforming many Vision-Language-Action (VLA) models, including OpenAI’s ChatGPT-4o and Anthropic’s Claude 3.5 Sonnet, in benchmark tests.
Who are the major players in the Physical AI race?
Several global tech companies are investing heavily in Physical AI:
- Google (DeepMind): Leading the way with its Gemini Robotics models.
- Nvidia: Introduced the Cosmos platform for Physical AI progress, predicting a “ChatGPT moment” for robotics.
- Microsoft: Released its VLA model “Magma”.
- Hugging Face and Physical Intelligence: Introduced “Pi0,” an open-source VLA model.
- UBTECH (China): Developed the Walker R1 bipedal robot, already integrated into production lines.
What is China’s strategy for Physical AI development?
China recognizes the transformative potential of Physical AI, referred to as “embodied intelligence” (å ·èº«æºè½). Specific actions include:
- Government Support: Inclusion of “embodied intelligence” in the government’s work report at the two Sessions.
- Policy Initiatives: Outlining the “2025-2027 Embodied Intelligence Technology Innovation and Industrial Development Plan.”
- Open-Source Platforms: Launching the world’s first general-purpose humanoid robot open-source platform by the Beijing Humanoid Robot Innovation Center.
- Industrial Integration: Integrating robots like UBTECH’s Walker R1 into manufacturing processes.
What are Vision-Language-Action (VLA) models?
Vision-Language-Action models are AI systems that combine visual perception (vision), natural language understanding, and the ability to perform physical actions. these models can “see” the world through cameras, “understand” instructions given in human language, and then execute those instructions through robotic movements and manipulation.
What is the importance of Nvidia’s Cosmos platform?
Nvidia’s Cosmos is a platform designed to accelerate the development of Physical AI. It provides tools and infrastructure for training,simulating,and deploying AI models for robotics and embodied intelligence. Jensen Huang, Nvidia’s CEO, believes that Cosmos will be a catalyst for a major breakthrough in robotics, similar to the impact of chatgpt on natural language processing.
What industries will Physical AI impact?
Physical AI has the potential to revolutionize numerous industries by enabling robots and bright systems to perform complex tasks autonomously. Some key sectors include:
- Manufacturing: Automating production lines, improving efficiency, and enhancing quality control.
- Logistics: Optimizing warehouse operations,streamlining delivery processes,and managing inventory.
- Healthcare:Assisting in surgery, providing elderly care, and dispensing medication.
- Agriculture: Monitoring crops, harvesting produce, and managing resources efficiently.
- Automotive: Developing self-driving vehicles and improving vehicle manufacturing processes.
Key Companies and Technologies in Physical AI
| Company/Technology | Description | Key Contributions |
|---|---|---|
| Google DeepMind (Gemini Robotics) | AI models for robotics using Gemini 2.0 | Versatile robots capable of understanding natural language and adapting to environments. |
| Nvidia (Cosmos) | Platform for Physical AI development | Provides tools and infrastructure for training and deploying robotics AI models. |
| Microsoft (Magma) | Vision-Language-action (VLA) model | Enables robots to perceive and interact with their environment based on language commands. |
| Hugging Face & Physical intelligence (Pi0) | Open-source VLA model | Transforms natural language commands into autonomous robot actions, fostering accessibility. |
| UBTECH (Walker R1) | Bipedal robot walker | Integrated into production lines, showcasing industrial applications of humanoid robots. |
