
Contributing a Force-Rich Reference Dataset to the Field
The BLO LAB reference dataset is a curated slice of our production capture, released with force, pose, and video aligned to the standards a serious tactile benchmark will eventually require. Here is what it contains, why we ship it, and what we hope the field does with it.
The commercial datasets we license to customers are not the whole picture of what a data supplier owes the community. Every serious infrastructure company in adjacent fields, compilers, databases, sensor hardware, contributes a reference artifact that lets the community evaluate the ecosystem independently of any single vendor. In robot data, that reference artifact is a curated dataset that anyone can download and train on. This piece describes ours.
What the reference dataset contains
A curated slice of our production capture across a small but representative task library: kitchen preparation, workshop assembly, laundry folding, and dexterous small-object handling. Each task is represented by hundreds of demonstrations from multiple operators, with all four capture streams, per-finger force, pose, first-person video, and, where relevant, whole-body suit data, aligned to sub-millisecond precision.
The tasks were chosen because they exercise the parts of the space that vision-only datasets under-represent. Contact-rich handling of deformable objects. Bimanual coordination. Slippery and fragile grasps. Fine-grained placement. The reference set is not a benchmark by itself, but it is a plausible seed for one.
What it is not
It is not our full production corpus. Customers who license our data receive volumes and task breadth well beyond what the reference set contains. The reference set is intentionally small enough to be usable on a single-GPU training run so that individual researchers and small teams can work with it, and intentionally representative enough that results on it correlate with results on the full corpus.
It is also not a leaderboard. We publish the data with documented protocols for a small set of evaluation tasks, but we do not host a leaderboard, do not rank submissions, and do not intend to. That is a role for a neutral community body, not for a supplier.
Why we ship it
Three reasons. First, the tactile side of the field is data-poor enough that a well-curated public reference set is genuinely useful. Researchers who cannot afford to license commercial data can still work on force-conditioned policies. The whole field benefits when more people can build on the same starting point.
Second, we want the assumptions our production data is built on to be inspectable. The alignment protocol, the calibration methodology, the labeling taxonomy, these are described in the reference dataset's documentation and can be audited by anyone. Trust in a commercial supplier compounds when the practices behind the paid product are visible in a public one.
Third, the reference dataset is a filter. Teams that engage seriously with it, publish results on it, and come to us for volume are qualified customers by definition. The reference dataset does customer discovery in a way marketing cannot.
What we hope the field does with it
We hope researchers use it to demonstrate that force-conditioned policies outperform vision-only baselines on the tasks where the difference should matter. We hope other data suppliers publish comparable reference sets so that the community can evaluate multiple sources on comparable terms. We hope the benchmark protocols that eventually consolidate around tactile learning are informed by the properties of these reference sets rather than by the properties of whichever narrow research setup happened to publish a paper first.
The larger hope is that the tactile side of manipulation research moves in the next two years to the position vision moved to fifteen years ago, a broadly-adopted reference dataset, a set of well-understood benchmarks on top of it, and a supply chain of commercial data that connects into the same infrastructure. That transition is what turns tactile learning from a specialization into a foundation.
The through-line
Across the eighteen pieces in this dimension, the through-line is straightforward. The tactile capabilities the robotics field needs are not model-limited; they are data-limited. The data is not scarce for lack of interest; it is scarce because building it at scale requires a specific stack of hardware, calibration, alignment, and operator infrastructure that few organizations have assembled. BLO LAB was built to be one of those organizations, and the reference dataset is the visible tip of the pipeline that produces the commercial data behind it.
The next dimension in this series turns to the operator economy that makes distributed capture work at all, and to the specific role Talika plays in making skilled human demonstration a real global market.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.


