Google DeepMind's Gemini Robotics 2 combines vision, language, and action models to control multiple types of robots, including humanoids performing tasks such as organizing shelves, tying bags, and replacing lightbulbs. "It's another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can," Carolina Parada, head of robotics at Google DeepMind, tells WIRED. From the report: Gemini Robotics 2 combines several different AI models into a single system. Taken together, they allow a robot to make sense of its surroundings and how to act in it. A vision language model (VLM), which understands images and video, can communicate with humans and reason how to perform different tasks. Two vision language action (VLA) models, trained to understand how to move in physical space, control the robot's full-body movement as well as the movements of grippers or hands. In video demonstrations shared ahead of the release, the company showed several different robots performing complex tasks autonomously using the amalgamated model. In one demo, Apptronik's Apollo 2 robot used hands from a company called Sharpa to tidy shelves. Google DeepMind trained the model to perform these tasks using a mix of human teleoperation, video examples, and simulations -- it's not yet possible for AI models to perform a wide range of complex tasks without specific training. [...] Parada says Google takes a multi-layered approach to safety, with guardrails applied on each model layer. It's also introducing ASIMOV-Agentic, a new benchmark for measuring the safety of various AI systems collaborating to control a robot. The benchmark detects whether a command will result in harmful or uncertain outcome.Read more of this story at Slashdot.