All articles
Case Study·11 min read·March 30, 2026

Kitchen Robotics: Why Cooking Is the Ultimate Benchmark

Kitchens are chaotic, deformable, and unforgiving. That is exactly why they are the right frontier.

Warehouses are neat. Factories are structured. Kitchens are neither. That is what makes kitchen tasks the honest benchmark for general-purpose manipulation, and why every serious humanoid company keeps circling back to cooking demos. If your robot can operate a kitchen, it can probably operate most of the built human world.

Every object is a special case

Onions slip. Herbs bruise. Dough sticks to gloves and to itself. Cast iron is heavy. Wooden spoons are compliant. Plastic cutting boards move when you press on them. Water is unpredictable at every scale, from a splash to a droplet. A robot that cooks has to negotiate all of these in a small, cluttered workspace with sharp edges.

This variety is exactly what forces policies to learn the underlying physics instead of memorizing task templates. A pick-and-place benchmark can be gamed with the right primitives. Cooking cannot.

Failure is visible and unforgiving

A dropped bolt in a factory is a minor recovery event. A dropped egg on the floor is a demo-ending event. Broken glass, spilled oil, burned food, these are irreversible failures that make cooking benchmarks brutally honest. Policies that look good on synthetic benchmarks often collapse the first time they meet a real crack of a shell.

This is a feature, not a bug. The gap between 'looks like it worked' and 'actually worked' is exactly where robot capability lives.

Contact is continuous, not discrete

A pick-and-place task has two contact events: grasp and release. A cooking task has continuous contact throughout, the knife on the board, the spoon in the pot, the hand on the whisk. Force control matters at every moment, not just at the beginning and end. This is what makes force data indispensable for the cooking frontier.

Our datasets

The kitchen slices of the BLO library, peeling, kneading, dish washing, dough handling, salad preparation, egg cracking, sauce stirring, are recorded with the full sensor stack: egocentric video, wrist video, whole-body IMU kinematics, per-finger force, and audio (the sound of a knife on a board is a surprisingly strong policy signal). Each task set contains hundreds of trajectories across multiple operators, multiple environments, and deliberate variation in the objects manipulated.

They are hard on purpose. A policy that succeeds on our dish-washing set is a policy ready to be deployed in a real kitchen. A policy that fails is a policy that has more to learn.

Why start here

There is a temptation to start with easier tasks and work up. The teams that ship do the opposite: they take on a maximally hard, maximally variable domain first, and let the difficulty force the pipeline improvements that would otherwise be endlessly deferred. Cooking is that forcing function.

Work with us

Building or training robots?

We license manipulation datasets and run custom capture programs. Get in touch to see what fits.