All articles
Industry·12 min read·July 1, 2026

The Teleoperation Ceiling: Why Warehouses of Robot Arms Cap Out

Robot-in-the-loop capture, an operator flying a leader arm that drives a follower arm, is the fastest way to start collecting manipulation data. It is also the mode that hits a wall first.

The default architecture for collecting manipulation data in 2026 is straightforward. Rent a warehouse. Fill it with robot cells. In each cell, put a follower arm holding the real end-effector and a leader arm, usually a low-cost puppet like GELLO, that the operator moves by hand. The follower mirrors the leader in real time. The operator watches through the robot's cameras and drives the task. Every session is logged.

It is a beautiful architecture. It produces demonstrations that are natively in the robot's action space, which means the policy trained on them has no retargeting problem. XDOF has raised $70M on a version of this thesis, and the frontier labs are buying. The pipeline works.

The pipeline also has a ceiling that is easy to see once you draw the throughput math on a whiteboard, and it is the reason no one who has thought seriously about the space believes teleoperation warehouses will be the only pillar of robotics data in five years.

The throughput math

A robot-in-the-loop cell is bounded by three fixed costs. The arm itself, a bimanual industrial-grade setup is $50–150k per station. The floor space, a real cell with a table, cameras, lighting, and safety cordoning is fifty to a hundred square feet. And the operator, one skilled human, one station, in-person, on a schedule.

None of these three scale sub-linearly. Doubling your operator hours means doubling your arms, your square footage, and your electricity bill. The unit economics improve slowly with volume, but they do not bend. This is the shape of a services business, not a software business, and every serious data-infrastructure company running teleop warehouses eventually confronts it.

The number that governs the whole industry is dollars per usable operator-hour of demonstration. In a teleop-warehouse model that number floors somewhere in the low-to-mid three figures once you fully load it. That is fine for the first million hours of frontier-lab data. It is not fine for the hundred million hours a general-purpose brain is eventually going to want.

The embodiment lock-in

The second ceiling is that every demonstration is captured on the specific robot embodiment in that cell. Change the arm, and the data does not transfer cleanly. Change the gripper, and it transfers even less cleanly. A warehouse configured for a two-finger parallel gripper cannot produce data for a five-finger dexterous hand without a full hardware refit.

This is fine for a lab targeting one deployment platform. It is a serious problem for anyone building a horizontal foundation model that has to run on many robots. The teleop warehouse is a single-embodiment machine, and building a second warehouse for a second embodiment is a full capex event.

Where the ceiling shows up in practice

The frontier labs already know this. It is why the largest of them are running teleop programs and wearable programs in parallel. Teleop for the tasks that need to be in the deployment action space today. Wearable for the tasks that need volume, embodiment diversity, and skilled operators who are not going to relocate to a Berkeley warehouse for a six-month contract.

The two modes are not competitors. They are complementary layers of the same stack. But if you can only afford to build one, and your bet is on a universal brain that has to work across embodiments, the layer that scales without a floor plan is the one that compounds.

Work with us

Building or training robots?

We license manipulation datasets and run custom capture programs. Get in touch to see what fits.