Google DeepMindhas introduced Gemini Robotics 2, a new model family for humanoid and other robots. On Apptronik’s Apollo 2, the AI shows continuous whole-body control for the first time. The robot can walk, bend down, grab objects and carry out complex tasks independently. At the same time, the system improves fine motor skills, the collaboration of multiple robots and operation without a permanent cloud connection. Google DeepMind wants to turn specialized machines into adaptable assistants that can understand their environment, plan actions and react to changes.
The most important thing in brief
- Gemini Robotics 2 connects visual perception and language directly to robot movements.
- Apptronik’s Apollo 2 can use it to coordinate its entire body from the foot to the fingertips.
- Three specialized models handle whole-body control, embodied thinking and local AI processing.
- Gemini Robotics ER 2 plans longer processes and enables different types of robots to work together.
- The ASIMOV-Agentic security benchmark checks, among other things, whether robots reject dangerous commands and request human help in a timely manner.
What is Gemini Robotics 2?Gemini Robotics 2 is a family of Google DeepMind models for autonomous robots. It processes images and voice commands, plans multi-step tasks and translates decisions into coordinated movements of the entire robot body.
Gemini Robotics 2 becomes the intelligence layer for robots
Google DeepMind calls Gemini Robotics 2 an “intelligence layer” for robots. The platform is intended to take machines beyond rigid and pre-programmed workflows. A robot should therefore not only repeat stored movements. Instead, he perceives his surroundings and interprets a linguistic order. He then decides which action steps are necessary for the desired goal. During execution it can react to changing conditions. This brings robots closer to systems that can handle tasks in less predictable environments.
The model family is not aimed exclusively athumanoid robots. According to Google DeepMind, it can also be transferred to double-arm systems and other types of robots. Adaptations to new onesHardwareshould sometimes be possible within a few hours. This is important because robots differ significantly in size, joint structure, sensors and gripping technology. A versatile AI platform could therefore reduce the effort required to develop your own control models. Manufacturers would then not have to train all the skills for each robot from scratch. The software would thus be more decoupled from a specific mechanical platform.
The launch builds on the ongoing collaboration between Google DeepMind and Apptronik. Apptronik opened the so-called Robot Park at the beginning of July 2026. At this facility, Apollo 2 robots perform practical tasks while collecting training data. This data will be incorporated into the development of the next generation of robotics models. This also includes Gemini Robotics 2. The Robot Park combines real robot work with continuous AI training. Apollo 2 serves both as a demonstration platform and as a source of new movement and interaction data.
Three AI models take on different tasks
Gemini Robotics 2 consists of three models, each with its own function. The main model combines perception, language and movement. Gemini Robotics ER 2, on the other hand, handles planning and embodied reasoning. The on-device version is designed for local operation on the robot hardware. Together, the models cover central levels of an autonomous robot. This includes understanding a task, planning the action and physically carrying it out. This allows Google DeepMind to use the models depending on the respective application.
| Model | Model type | Central function | Appropriate systems |
|---|---|---|---|
| Gemini Robotics 2 | Vision-Language-Action model | Translates visual and verbal input into movements and controls the entire body | Humanoid robots and double-arm systems |
| Gemini Robotics ER 2 | Embodied reasoning model | Plans multi-step tasks, communicates with humans, and coordinates multiple robots | Single robots and mixed robot fleets |
| Gemini Robotics On Device 2 | Locally running VLA model | Works directly on the hardware and can be adjusted with just a few training data | Robots with limited or unreliable cloud connection |
The abbreviationVLAstands for “Vision-Language-Action”. Such a model combines images, verbal instructions and concrete actions. For example, it recognizes an object and understands its position in space. It then converts an order into appropriate movement commands. The robot does not have to receive each individual step as a separate instruction. Rather, a command can describe a complete goal. This direct connection makes more natural human-robot interactions possible.
Gemini Robotics ER 2 complements this level of execution with longer thinking and planning processes. The model can decompose multi-stage tasks and monitor their progress. It also communicates with humans and other robots. Gemini Robotics On-Device 2, on the other hand, focuses on efficient local execution. This creates a model family that combines perception, planning and movement. However, the three names do not necessarily mean three systems isolated from each other. In a practical application, their skills can intertwine.
Apollo 2 will receive full, whole-body controls
The most striking innovation is the full-body control of Apptronik’s humanoid Apollo 2. Previous versions of the robotic models focused primarily on manipulations with the upper body. Gemini Robotics 2 now coordinates movements from the feet to the fingertips. Apollo 2 can therefore walk, bend and tilt its upper body autonomously. At the same time, he can recognize objects, grab them and move them to another location. The system must continuously monitor the balance of the humanoid robot. Perception, movement and manipulation thus become part of a common process.
Google DeepMind demonstrated this capability with a concrete voice command. Apollo 2 was supposed to place a watering can in a green container on the bottom shelf. The robot first had to understand the order and the objects mentioned. He then walked across the room to the watering can. He picked up the item and transported it to the shelf. Because the target container was in the lower area, Apollo 2 had to adjust its posture accordingly. Finally, he placed the watering can in the desired location.
This demonstration initially seems like a simple transportation task. However, it requires the coordination of numerous sub-skills. The robot must visually distinguish the watering can and the green container. He also needs a spatial idea of the room and a safe walking route. When grasping, he must take the shape and position of the object into account. He must not lose his balance or control of the watering can during transport. He also has to recognize the height of the target area on the shelf. Only the interaction of these steps turns a single movement into an autonomous whole-body action.
Hands, grippers and industrial precision tasks
In addition to locomotion, Google DeepMind has improved the robots’ dexterity. Apollo 2 can be equipped with five-joint SharpaWave robot hands. Gemini Robotics 2 uses it to control complex movements of individual fingers. In a demonstration, the robot tied a garbage bag. In another task he unscrewed a light bulb. Both activities require controlled movements and ongoing adjustment of the force used. They therefore go well beyond simply picking up a solid object.
Tying a garbage bag is particularly challenging for robots. Plastic film is soft, malleable and changes shape with every touch. The robot cannot therefore record the exact position of the material once and then act blindly. He must correct his movements based on new visual information. When unscrewing a light bulb, other difficulties arise. The hand must grasp the object securely and make a rotational movement. At the same time, she must neither drop the light bulb nor damage it with too much force.
However, Gemini Robotics 2 doesn’t just support human-like hands. The model can also control conventional parallel grippers. Such grippers are widely used in industrial plants. They are suitable for repeatable and precise handling tasks. Google DeepMind mentions the precise insertion of components as possible applications. Packaging and sorting work is also an option. The platform combines the versatility of humanoid hands with the robustness of industrial gripping tools.
| Gripping technique | Task shown or mentioned | Special requirement |
|---|---|---|
| Five-fingered SharpaWave hands | Tie up garbage bags | Handling soft and malleable material |
| Five-fingered SharpaWave hands | Unscrew the light bulb | Precise force control and coordinated rotational movement |
| Parallel gripper | Precision assembly and insertion | Accurate positioning of components |
| Parallel gripper | Packing work | Reliable gripping and placing in recurring processes |
Several robots plan and work together
Gemini Robotics ER 2 is also intended to enable multiple robots to work together. Different types of robots can exchange information and coordinate their actions. A complex workflow no longer needs to be carried out entirely by a single machine. Instead, the system can distribute individual steps to suitable robots. A mobile robot could, for example, deliver materials. A humanoid robot could then take on an unstructured handling task. A stationary double-arm robot could then carry out a precise assembly step.
This cooperation is particularly interesting for factories and logistics centers. Various robot systems with specialized capabilities are often used there. Until now, they have often worked in separate areas or according to fixed handover rules. Gemini Robotics ER 2 could connect such systems more flexibly. To do this, the model must track the state of the entire workflow. It also needs to recognize when a task is completed or a handoff is required. Coordination thus becomes an independent AI task.
The reasoning model is designed for processes that last several minutes. Meanwhile, it monitors progress. If an action step fails, the system should not terminate immediately. Instead, it can detect the error and plan a new attempt. It can also adapt its strategy to changing situations. This skill is critical for real-world work environments. Objects are not always in exactly the same position, and people can also change a planned process.
Communication with people is also part of Gemini Robotics ER 2. The model can receive instructions and process queries or status information. This allows employees to guide a robot at the level of the work goal. You don’t have to program every joint movement or waypoint. However, it remains crucial how reliably the model recognizes ambiguous instructions. Responsibility for errors must also be clearly regulated. Clearly defined processes and safety limits are therefore still required for industrial applications.
Local execution and new security mechanisms
Gemini Robotics On-Device 2 runs directly on a robot’s hardware. The model is therefore not permanently dependent on a cloud connection. This feature is suitable for factories, warehouses and remote locations with limited network coverage. Additionally, local processing can reduce the delay between perception and movement. Sensitive camera and operating data do not have to be transferred to external servers for every decision. However, performance and responsiveness depend on the available hardware. Google DeepMind has not provided general performance metrics for every conceivable robotics platform.
The low adjustment effort is particularly relevant. According to Google DeepMind, On-Device 2 can be transferred to new dual-arm robots. Less than 200 training examples should be sufficient for this. The data can be collected within a few hours. This could significantly speed up the commissioning of new robots. At the same time, it remains unclear how much the amount of data required will increase with particularly complex hardware or unusual tasks. The threshold mentioned is therefore to be understood as the manufacturer’s information for the demonstrated adaptation scenarios.
Google DeepMind adds the security benchmark ASIMOV-Agentic to the model family. He examines agentic security coordination and dealing with uncertainty. For example, it is checked whether a robot refuses a dangerous order. It should also recognize when a task cannot be carried out safely. In such a case, the system must request human assistance. The benchmark therefore views safety as more than just an emergency stop function. It also checks whether the AI appropriately assesses risks during planning.
According to Google DeepMind, Gemini Robotics ER 2 also has improved human perception. The system is intended to detect people nearby and trigger appropriate security functions. If a human enters the work area, the robot can stop in a controlled manner. This capability is particularly important when autonomous machines do not work permanently behind protective fences. However, such AI functions do not replace mandatory industrial security concepts. Sensor technology, risk assessments and independent protection systems remain necessary for real operations. The functions presented therefore form an additional layer of security.
New perspective:The most important change may lie less in a single spectacular movement than in the shortened learning time. When an on-device model can actually be adapted to new hardware with fewer than 200 examples, the economics of small automation projects change. Robots could then also take on tasks that were previously too expensive to program traditionally. At the same time, however, a new dependence on the quality of a few training examples arises. Incorrect or biased data could translate more quickly to real movements. Companies therefore not only need robot technicians, but also procedures for checking training data, behavioral limits and updates.
Gemini Robotics ER 2 is already available via Google AI Studio. The model is also in private preview on the Gemini Enterprise Agent Platform. Access to the other two models is more restricted. Gemini Robotics 2 and Gemini Robotics On-Device 2 are initially available to select early access partners. A general release was not mentioned in the announcement. There is also no information on prices or specific commercial deployment dates. The capabilities demonstrated therefore mark an important development step, but not yet a comprehensive market launch.
Conclusion
Gemini Robotics 2 shows how quickly humanoid robots transform from programmed tools to adaptable action systems. Apollo 2 runs, attacks and plans as a coordinated unit for the first time. What is particularly exciting, however, is the local processing, the quick adaptation to new hardware and the cooperation of several machines. Whether this results in reliable factory helpers is determined by real endurance tests, safety evidence and operating costs. Google DeepMind provides a powerful intelligence layer for this. Now it has to prove, outside of controlled demonstrations, how robust, safe and economical it really works.
Classification: What the demonstration shows – and what it doesn’t
Controlling a humanoid robot through a robotic model is an important step for embodied AI. However, a demonstration does not automatically prove robust series operation. Repeatability, safety limits, error handling, speed and performance in changing real-world environments are critical.
FAQ
What is Gemini Robotics?
Gemini Robotics is an AI approach from Google DeepMind focused on robotic perception, planning and action. It is intended to connect language and perceptual information with physical actions.
Does this mean Apollo 2 can operate completely autonomously?
A control or demo shown does not automatically mean complete autonomy in every environment. The specific level of use depends on the task, sensors, safety concept and training.
Why are humanoid robots interesting for factories?
You can work in environments that are already designed for people, tools and shelves. They only become economically relevant when performance is reliable, repeatable and maintainable.
Sources
- Google DeepMind: Gemini Robotics
- Apptronics: Apollo
Author Nico Nuss has been working on mobile computing and automation software since 2001. Drawing on his experience and strong interest in future technologies, he focuses on robotics and AI.
![[Image content created with AI] cropped ALPHA BIONIC LOGO [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2025/12/cropped-ALPHA-BIONIC-LOGO.png)
![Gemini Robotics Controls Apollo: What the Humanoid Demo Means 1 [Image content created with AI] Gemini Robotics 2 [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/08/Gemini-Robotics-2.png)
![Gemini Robotics Controls Apollo: What the Humanoid Demo Means 2 [Image content created with AI] Nico Nuss [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2025/12/Nico-Nuss_1-150x150.jpg)
![Humanoid Robots Under 25000 Euro: Prices, Limits and Buying Guide 3 [Image content created with AI] KI-generierte Symbolaufnahme: Frau mit humanoidem Roboter in einer Wohnung [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/08/humanoide-roboter-unter-25000-16x9-1.png)
![URKL: The fighting league for humanoid robots 4 [Image content created with AI] URKL: Die Kampfliga für humanoide Roboter [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/07/URKL-Roboter-Liga.png)
![EU AI Act: Everything companies need to know now 5 [Image content created with AI] EU AI Act [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/04/EU-AI-Act.png)
![The best AI image generators 2026: Create images online for free 6 [Image content created with AI] Die besten KI-Bildgeneratoren 2026: Kostenlos online [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/07/KI-Bildgeneratoren-kostenlos-online.png)
![Promptchan AI: features, costs and risks 7 [Image content created with AI] Promptchan AI [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/07/Promptchan-AI.png)
![Remove AI watermark: 4 easy methods 8 [Image content created with AI] KI Wasserzeichen entfernen [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/07/KI-Wasserzeichen-entfernen.png)
![Cancel ChatGPT subscription: Instructions for Web, iPhone & Android 9 [Image content created with AI] ChatGPT Abo kündigen [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/07/ChatGPT-Abo-kuendigen.png)
![Figure 03: What the humanoid robot can really do 10 [Image content created with AI] Figure 03 – humanoider Roboter für Haushalt und Gewerbe [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2025/12/Figure-03.png)
![Humanoid Robots and the AI Singularity 11 [Image content created with AI] humanoide roboter ki singularitaet 16x9 1 [Image content created with AI]](https://alpha-bionic.info/wp-content/uploads/2026/07/humanoide-roboter-ki-singularitaet-16x9-1.png)