
The 90% of Manipulation Datasets That Skip Force Entirely
A frank audit of the public and semi-public manipulation datasets that dominate foundation-model training. Force appears in a small minority, is calibrated in a smaller minority, and is present at usable rates in a smaller minority still. The gap is a strategic opening.
It is worth being specific about how underrepresented force actually is in the datasets that train modern robot policies. Rough impressions circulate, 'nobody has force', and rough impressions are unhelpful. The picture becomes actionable when you count.
What the count looks like
Take the ten largest public manipulation corpora used in foundation-model training. Open X-Embodiment. Bridge V2. RH20T. RoboSet. DROID. FMB. TACO-Play. And several proprietary collections that have leaked enough documentation to be counted honestly. Aggregate trajectories across these sources easily exceed a million.
The fraction of trajectories that include any force or torque signal is somewhere between three and eight percent, depending on how permissively you count. The fraction that include a fingertip force signal, not a wrist wrench, an actual per-finger force reading, is under one percent. The fraction that include a per-finger force reading at rates sufficient to capture slip transients is essentially unmeasurable because there are not enough examples to form a statistic.
This is not a statement about research quality. There are excellent tactile datasets, Feeling of Success, the various DIGIT collections, ContactDB, that focus on force at the expense of scale. It is a statement about what the large training runs are actually seeing. When a foundation model is trained on ninety-something percent vision-and-pose data, the resulting policy is a vision-and-pose policy, with everything that implies.
Why the shortage is structural
Force sensing has been treated as a specialization inside manipulation research rather than a baseline capability of a capture rig. Labs that focus on tactile build careful, small, well-instrumented datasets. Labs that focus on scale build large, sparse, force-free datasets. The two lines rarely cross, because the operational demands are different: force-focused capture wants controlled conditions and calibrated sensors; scale-focused capture wants many operators, many locations, and minimal setup overhead.
The consequence is that the datasets which foundation-model teams reach for at training time are almost entirely on the scale-first side of the split. Not because those teams do not value force, but because there is not yet a body of force-rich, well-calibrated, license-clean data at the scale foundation-model training requires.
What happens because of it
Foundation models trained on this mix inherit predictable weaknesses. They handle rigid, well-lit objects well and deformable, cluttered, or novel objects poorly. They struggle with tasks whose failure modes are force-defined, insertion, pouring, cloth manipulation, bimanual handover, even when the visual scene is trivially interpretable.
The gap shows up most starkly in benchmark reports. Policies trained on the standard mixtures post strong numbers on pick-and-place and weak numbers on contact-rich tasks, and the ratio between the two is remarkably stable across model families. That stability is the fingerprint of a data problem, not a modeling problem. When the same weakness recurs across architectures, the shared input distribution is doing the work.
Where the opening is
The opening is straightforward. A team that produces manipulation data at scale, with per-finger force calibrated across operators and sessions, at rates that preserve slip transients, in formats and licenses that foundation-model teams can consume, occupies a category with essentially no competitors. Not because the technology is out of reach, it is not, but because building it requires bridging the tactile-research culture and the scale-capture culture, and few organizations have done both.
This is the category BLO LAB was built to fill. The glove exists because the counting exercise above pointed at a specific hole in the training-data supply, and closing it required a capture instrument designed against that specific shortage rather than against the demo-friendly incentives of a research lab.
The next piece describes the signal chain that makes calibrated per-finger force at scale actually work.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.


