← All posts

Gemini Robotics 2 can tie a knot. The dustpan beats it

2026-08-04 · 4 min read

On July 30, Google DeepMind launched Gemini Robotics 2. The post carrying it was five words: @GoogleDeepMind opened with "One brain. For any robot."

Demis Hassabis, who runs the lab, posted that robots can now "reason through every movement," pointing at knot tying and at different machines teaming up on jobs a single one could not finish. The clips are genuinely impressive. The success rates published alongside them are the part worth reading first.

What actually shipped

There are three models, and they split the job the way a person does.

Gemini Robotics 2 turns a camera feed and a spoken instruction into motor commands, now across a whole body instead of arms on a tabletop. It walks, crouches, and reaches for things on the floor. ER 2 is the planner. It handles tasks that run several minutes and hundreds of decisions, notices when a step has gone wrong, and coordinates more than one robot. On-Device 2 runs locally with no network round trip.

The claim underneath "one brain, any robot" is the interesting one. A single model checkpoint drove Apptronik's Apollo 2 humanoid with SharpaWave hands, the same Apollo 2 wearing different hands from Inspire, and a Franka Duo with a plain two-finger gripper. Three bodies, one set of weights.

@kimmonismus flagged the number most coverage skipped: the on-device model can "adapt to a completely new two-arm robot with fewer than 200 examples." Robot learning has historically wanted thousands of demonstrations per platform, per task. Under 200 is a different category of effort.

The numbers in the fine print

DeepMind published per-task success rates, and they do not match the vibe of the launch video.

Whole-body picking with Apollo 2 and Inspire hands: 76.3 percent off a shelf, 68.4 percent off a table, 45.7 percent off the floor. Picking something up from the ground fails more often than it works.

Five-fingered dexterity spreads much wider. Unscrewing a lightbulb, 92 percent. Screwing one back in, 36 percent. Tying a trash bag, 44 percent. A ziplock, 40 percent. Sweeping with a dustpan, 32 percent.

The knot tying is real. So is a two-in-three failure rate on a chore a nine-year-old picks up in an afternoon. DeepMind says as much in its own materials, noting that movement speed still has a way to go and that the on-device model struggles with unfamiliar tasks and high-degree-of-freedom hardware.

The honest read

This is a real advance and a demo at the same time, and both readings are correct.

The advance is generalization. For most of the history of robotics, a system was welded to one body and one task, and moving it anywhere else meant starting over. A checkpoint that survives a hand swap and a hardware swap is the thing that changes the economics, because it means the training investment stops evaporating every time the hardware changes.

The demo part is everything else. ER 2, VLA, and the on-device model are behind an early-access program and a private preview. Nobody outside a trusted tester list can check any of this. And a 32 percent success rate is not a product. It is a promising research result that will look silly in eighteen months, which is exactly what happened to the robot video reels from 2023.

What is not in question is the direction. We wrote about this same pattern nine days ago when a video generation lab found its model could drive factory robots. Two different labs, two different starting points, converging on general models that carry skill across bodies.

What this signals if you run a business

Neither takeaway requires you to care about humanoids.

  • Ask for the failure rate before you fall for the demo. Every AI tool sold to you this year will be pitched with its lightbulb number, not its dustpan number. The question that matters is how often it misses, on your messiest inputs, and what happens next when it does. A system that works 76 percent of the time and tells you when it failed is worth more than one that works 90 percent of the time silently.
  • Teaching a system a new task is getting cheap. Under 200 examples to learn a new robot body is the same trend that lets a quoting or intake workflow learn how your shop actually talks after a couple dozen real jobs. The setup cost that made custom automation a big-company thing keeps falling.

That second point is the one with money attached. Work that was not worth automating two years ago, because the configuration took longer than doing it by hand, is worth a second look now. If you want an outside read on which parts of your week qualify, New Face Design does a free process audit for businesses in the Fox Valley and anywhere else.

A humanoid tied a knot on video last week and failed the dustpan two tries out of three. Both facts belong in the same sentence, and most of the coverage only printed one of them.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere