
The GELLO Legacy: What Comes After the $300 Puppet Arm
The GELLO teleoperation rig lowered the cost of imitation-learning data collection by an order of magnitude. Its own authors are now building the company that has to answer what comes next.
In 2023, a Berkeley paper introduced GELLO, a general, low-cost, intuitive teleoperation framework for robot manipulators. The core insight was that a kinematically similar puppet arm, built from off-the-shelf servos for under $300, was a dramatically better human interface than a VR controller or a space-mouse. The paper became one of the most-cited teleoperation designs of the following two years, and it seeded a generation of open-source imitation-learning setups.
The authors of that paper are now the founders of XDOF, a well-funded data-infrastructure company built on the industrial version of the same idea. GELLO worked. The question the industry is now asking is what comes after it.
What GELLO solved and what it did not
GELLO solved the operator interface. A skilled human could drive a robot arm with GELLO within a few minutes of first touching it, produce trajectories smooth enough to train on, and do it for hours without the physical strain a VR controller inflicts. That was a genuine step change in usability, and it is why the design became the default.
GELLO did not solve, and was not intended to solve, the deeper problems of manipulation data at scale. It did not add force or tactile channels, the leader arm is unpowered, so there is no proprioceptive force to record. It did not remove the need for a follower arm and its cell. It did not address the embodiment-lock-in problem. It made the existing pipeline better; it did not replace it.
What the successor has to answer
Any credible successor to the GELLO paradigm has to be honest about which of its limits are fundamental and which are engineering. The cost-per-hour ceiling of teleop warehouses is fundamental, it comes from having a physical robot in the loop. The lack of force signal is engineering, and instrumenting a leader arm to record simulated force helps but does not match a real fingertip sensor. The embodiment lock-in is fundamental to the architecture.
The industry response has split into two camps. One camp, XDOF's, is to industrialize the teleop warehouse and drive its cost down through operations, tooling, and scale. The other camp is to skip the robot entirely and capture directly from the human, using gloves, suits, and straps that record what the person is doing and let the retargeting happen offline. Both have merit. Both are being funded. The one that scales fastest per dollar is likely to be the one that removes the most physical infrastructure from the collection loop.
Where BLO LAB sits in this lineage
We are on the wearable side of the split, deliberately. The GELLO paper is part of the intellectual context we build against, it is the strongest version of the argument that a low-cost teleoperation rig is the answer, and we think that argument runs out at exactly the scale the industry is now trying to reach. Direct human capture is the mode that scales past that ceiling. Both approaches will exist for years. The question is which one the marginal hour of frontier-lab training data will come from in 2028.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.


