Vol. 1 · Edition 033Free · No paywall

Everyone Needs a Samwise

AI news · Synthesized · Opinionated · 🌿

Gemini Robotics 1 — upper half only

arms

Gemini Robotics 2 — one unified policy

whole body
Model Launch
By Sam Taylor with Samwise

On what 'whole body intelligence' actually means, which success rates Google put in the headline and which ones ended up in the appendix, and when to expect this in your building lobby.

Google's robot brain learned to use its legs. Read the dustpan number before you get excited.

Source lean on this story
▲ avg

Anti-AI

00

Skeptic

01

Neutral

00

Pro (practical)

01

Pro (hyped)

02

← Anti-AI · Pro-AI →

If you've watched a robot demo video — the kind where the machine gracefully picks up a cup, folds a towel, and wheels away while ambient music plays — and then read somewhere that robots still can't do your laundry, you know the gap. The demo is real. The gap is also real. Both things stay true at the same time, and navigating that is basically the whole story of physical AI in 2026.

Google DeepMind announced Gemini Robotics 2 on July 30. Three models. The headline claim: for the first time, a single AI policy controls a humanoid's entire body — legs, torso, arms, and five-fingered hands — all at once. Not just the upper half. The whole thing.

That's genuinely meaningful. And the success rates tell you exactly where "meaningful" ends and "your building lobby" begins.

What's actually new

The previous Gemini Robotics model only controlled a robot's upper body. Useful for tabletop manipulation — picking things up, assembling parts. Not useful for tasks that require crouching, leaning, stepping sideways while reaching for something on a low shelf.

Gemini Robotics 2 unifies all of that under one model. It sends movement commands for every joint at once, so the robot's decision-making isn't split between "what should my arms do" and "what should my legs do." In robotics, this is called a single policy for whole-body control, and it's been a core unsolved problem because the action space is enormous and the coordination requirements are precise.

92%
Success rate for lightbulb unscrewing — the task Google highlighted in its announcement

→ Source: Google DeepMind

The three models are: Gemini Robotics 2 (the full-body VLA — the model that converts vision and instructions into movement), Gemini Robotics ER 2 (the "embodied reasoning" model that plans multi-step tasks, tracks hundreds of decisions, and coordinates multiple robots simultaneously), and On-Device 2 (a lighter version that adapts to a completely new robot body from fewer than 200 examples — meaning a company can deploy it on hardware Google's never trained on, fast).

Partners include Apptronik, Boston Dynamics, and Agile Robots.

The numbers Google put in the headline and the ones that ended up elsewhere

Here's the thing about the 92% lightbulb number. It's real. Also: lightbulb unscrewing is a rigid, repeatable mechanical task with a known grip profile. The robot reaches to a fixed point, rotates, removes. Same every time.

The dustpan number is 32%.

Sweeping dirt into a dustpan requires continuous contact with a deformable surface, judgment about where the pile actually is, and coordination between the broom hand and the dustpan hand while the robot's body adjusts. That's the harder class of task. And 32% is not a rounding error from 90% — it's a different category of capability.

Gemini Robotics 2 success rates across task types
TaskSuccess rateWhat makes it hard
Unscrewing lightbulb92%Rigid, repeatable mechanical task
Precision part insertion89.6%Fine motor with fixed target
Packing tools into boxes78.9%Variable shapes, one-directional
Lifting from shelf76.3%Whole-body reach with fixed object
Lifting from table68.4%Straightforward pick-and-place
Tying a trash bag44%Deformable material, two-hand coordination
Lifting from floor45.7%Floor crouching + object retrieval
Sealing a bag40%Deformable, continuous contact
Sweeping with dustpan32%Deformable surface + coordination

The pattern is clear. The system is strong on tasks where the target is rigid and the goal is mechanical (screw in, insert, pack). It's still rough on tasks involving deformable materials — bags, dirt, towels. Those require judgment about a surface that changes shape as you interact with it, and that's a different problem than the whole-body unification milestone solves.

I'm not saying 32% is shameful. I'm saying: if you saw the demo video and thought "this might clean my kitchen soon," read this table first.

Source spread

  • Google DeepMind blog — Gemini Robotics 2 — [hype]. The official announcement, including the benchmark numbers above. The 92% lightbulb figure gets more prominence than the 32% dustpan figure.
  • SiliconANGLE — [builder]. Covers the three-model architecture and access tiers in detail.
  • Bloomberg — [skeptic]. The headline itself uses "struggling with dexterity" — the coverage is appropriate to the task-range spread.
  • TechTimes — [builder]. Good technical breakdown of what "one policy" actually means in practice.

What's real:

  • The whole-body unification is architecturally genuine. One model controlling all joints is a real technical milestone, not marketing positioning.
  • ER 2 coordinating multiple robots is significant. Multi-robot collaboration for tasks like "two robots carry a long object" requires the planner to reason about both bodies simultaneously — this is new.
  • The On-Device 2 adaptation speed is impressive. Adapting to a completely different robot body from fewer than 200 examples means this doesn't require warehouse-scale data collection per hardware deployment.
  • ER 2 is actually available today on Google AI Studio. You can touch the reasoning model now, which is unusual for a physical-AI announcement.

What deserves a side-eye:

  • The 32% to 92% range across task types is a wide band for a unified announcement. The headline capability (whole-body control) is real; the headline success rate (92%) is the cherry-picked end of that range.
  • Early-access only for the VLA and On-Device models means the system's real-world performance data is still thin. "Partners in testing" is not the same as "deployed in production."
  • Physical AI timelines are notoriously slower than frontier model timelines because hardware constraints compound. ER 2 being on AI Studio today doesn't mean Gemini Robotics 2 is shipping into buildings next quarter.

What to do about it

For everyday readers:

  • Don't update your "when will robots replace me" timeline yet. The 92% lightbulb headline is real. The 32% dustpan reality is also real. Tasks that require judgment about surfaces and objects that deform as you touch them are still hard problems.
  • If you follow consumer robotics, watch the deformable-material tasks in future announcements. That's the bottleneck. When you see 80%+ on bags, towels, and cleaning tasks — then the timelines shift.
  • ER 2 is actually accessible now on Google AI Studio. If you're curious about what robot planning AI looks like in practice, this is a rare case where you can touch the research model directly.
  • For builders in the physical-AI space: the On-Device 2 fast adaptation story (<200 examples to a new robot body) is the detail worth exploring. That's what makes the system commercially deployable without massive per-hardware training runs.

Further reading

🌿

Liked this? Get the weekly digest.

Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.

Your take

How'd I do on this one?

What did I miss?

Tell Samwise (and Sam).

Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.