
Should a Foundation-Model Lab Buy Data or Build It?
Every lab building a physical-AI foundation model faces the same make-or-buy question on data. The right answer depends on the shape of the model, not just the size of the budget.
A frontier robotics lab in 2026 has roughly three options for where its manipulation data comes from. Build the collection pipeline in-house. Buy dataset hours from a specialist like XDOF. Or run a hybrid, where some tasks are captured internally and the long tail is outsourced.
The right choice is not obvious, and the trade-offs are not primarily financial. They are strategic.
The case for building in-house
The data pipeline is the model. If you outsource it entirely, you outsource one of the most important variables in your model's quality. In-house capture lets you iterate on the task spec quickly, catch data-quality issues early, and keep the loop tight between what the model needs and what shows up in the training set.
For labs whose model roadmap is tightly coupled to specific tasks, a cooking model, a lab-automation model, the argument for in-house is very strong. The data spec changes weekly. An outside vendor cannot keep up.
The case for buying
Building a capture program is a full company inside your company. Hardware team, operations team, operator recruiting, payment infrastructure, QA pipeline. That is a distraction from building the model itself, and for most labs it is not a distraction the CEO wanted when they founded the company.
A vendor that has already solved these problems can produce dataset hours faster, cheaper, and more reliably than an in-house program that is still standing up its first cell or onboarding its first cohort of operators. This is the pitch, and for a large slice of the market it is the right one.
The hybrid that most serious labs land on
In practice, the labs furthest along run both. In-house capture for the tasks closest to the model roadmap, where iteration speed matters more than throughput. Outsourced capture for the volume, embodiment diversity, and skill breadth that no single lab can build in-house at the required scale.
The vendors that win the outsourced slice are the ones that treat themselves as an extension of the lab's model team, deep integration on task specs, shared metrics on model performance, and a data delivery format that plugs directly into the lab's training pipeline. This is the shape of the relationship that both teleop-based and wearable-based data vendors have to grow into.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.



