
Why We Chose Force-Instrumented Gloves Over Vision-Tactile Fingertips
The BLO LAB glove was designed against a specific trade-off: fingertip vision-tactile sensors are richer per contact, but instrumented gloves scale to the volume, embodiment, and task breadth a foundation-model dataset actually requires. Here is the reasoning, in full.
When we started designing the BLO LAB glove we spent months looking at whether to integrate a vision-tactile fingertip into the palm-side of each finger. GelSight fits on a robot end effector, and in principle it fits on a wearable, and if it did the resulting glove would have per-contact-patch signal richness that no strain-gauge or capacitive design would ever match.
We chose not to. The reasoning has held up as the field has moved, and it is worth writing down honestly, because the trade-off is real and does not go the same way for every use case.
What we were optimizing for
The BLO LAB glove exists to convert an hour of skilled human work into an hour of usable robot training data. That framing pins down the design goals: high per-hour data yield, low operator friction, embodiment agnostic, delivered in formats a foundation-model team can consume without a science project.
Every design choice traces back to those goals. If a sensor takes ten minutes to calibrate per session, it costs half an operator hour per shift. If it requires a fume hood to replace, it does not deploy in a distributed operator network. If it produces frames at video rate, it misses the corrections a human hand makes in the two hundred milliseconds around a contact event. If its geometry pushes the palm surface a centimeter off the object, the recorded contact is not the contact a bare hand would have made.
Vision-tactile fingertips are excellent in every metric except these four, and these four are what matter for capture throughput.
The specific tradeoffs, one by one
Throughput. A vision-tactile fingertip needs periodic calibration, elastomer replacement, and camera exposure tuning. A capacitive-and-piezoresistive glove needs a warm-up pass and a per-operator size fit. Across a shift, the glove costs minutes of overhead; the vision-tactile rig costs hours. Across an operator network delivering tens of thousands of hours per quarter, that ratio compounds into the difference between shipping and not.
Bandwidth. Force events in skilled manipulation contain energy up to several hundred hertz. Vision-tactile fingertips sample at video rates and physically cannot resolve the transients. The BLO LAB glove samples force at rates that preserve those transients, because the sub-100-millisecond corrections around a slip event are the training signal that separates a policy that can hold a knife from one that can chop.
Embodiment fidelity. A vision-tactile fingertip changes the geometry of the contact. The operator is no longer manipulating with their finger; they are manipulating with a sensor housing. The captured trajectory is off by the housing's thickness and stiffness, and any policy trained on it inherits that shift. An instrumented textile glove is thin enough that the operator's proprioception, grip choice, and micro-corrections remain the operator's own.
Retargeting. Data captured on a human hand has to be retargeted to whatever end effector the customer's robot uses. Retargeting works better when the source data represents genuine human dexterity, not human dexterity as filtered through a sensor housing. This is a subtle point that becomes obvious the first time a customer runs the same trajectory retargeted to two different grippers and one of them fails at the contact events the housing had smeared out.
What we gave up
We gave up per-frame contact-patch geometry. A GelSight fingertip would tell us the shape of the pressure distribution across the last square centimeter of contact. Our glove tells us the aggregate force and pressure at each fingertip and along the palm, at high rate, but not the fine-grained shape of that pressure. For tasks where the fine-grained shape is decisive, coin sliding, edge following, printed-texture identification, a specialized rig with vision-tactile fingertips is still the right instrument.
For the very broad class of tasks that dominate industrial and household work, grasping, insertion, deformable manipulation, bimanual coordination, tool use, the aggregate force signal plus high-rate pose plus first-person video is enough to train policies that behave. The evidence for this shows up in the customers who license our datasets and then run credible policies on hardware whose fingertips do not resemble ours at all.
What we would add later
The obvious next axis is a hybrid: instrumented glove for volume capture across a wide task library, augmented by a small vision-tactile rig for a subset of contact-critical tasks that receive extra depth. This is a division of labor rather than a replacement, and it is where our roadmap is pointing. The glove remains the primary instrument because the primary instrument in a foundation-model pipeline is the one that produces the most hours of usable data per week. Everything else is supplementary.
The lesson, if there is one, is that the right sensor for the demo is not always the right sensor for the dataset. Vision-tactile fingertips are the right sensor for the demo. Instrumented gloves are the right sensor for the dataset. The field will keep both.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.


