
What a Serious Tactile Benchmark Would Look Like
The field has vision benchmarks, manipulation benchmarks, and mobility benchmarks. It does not have a serious tactile benchmark. A design sketch for what one would need to include, and why building it is more useful than another leaderboard.
Benchmarks are how a research community coordinates its attention. They are also how a research community lies to itself, when the benchmark measures something adjacent to what actually matters. Tactile learning has both problems at once: no benchmark of adequate scope exists, and the smaller benchmarks that do exist have started to shape research choices in ways that reward benchmark performance more than deployed capability.
What a serious tactile benchmark would have to measure
First, contact behavior on tasks where contact is decisive. Insertion, cable routing, cloth manipulation, deformable object handling, tool use. Not pick-and-place with a tactile fingertip attached, that is a vision task with an unused sensor. Tasks where the correct action at the moment of contact is not derivable from the visual scene alone.
Second, evaluation under sensor variation. A benchmark that runs on a single fingertip design produces winners whose policies are tuned to that design. A benchmark that requires the same policy to run on multiple tactile modalities, vision-tactile, capacitive, piezoresistive, glove-integrated, produces policies that generalize across the sensor mix the field actually deploys.
Third, evaluation across skill levels. Success rate on a well-executed grasp is a poor measure because policies that overfit to nominal conditions score well and then fail in deployment. Success rate under perturbations, novel objects, and unexpected contact events is what deployment cares about.
Fourth, evaluation with real hardware and real objects. Simulated tactile benchmarks have their place and are useful for pretraining, but a benchmark that never touches a real object misses the physical variance that makes tactile hard.
Why current tactile benchmarks fall short
Most current tactile benchmarks were built by individual labs to support the sensor design or the specific model architecture they were developing. This is not a criticism, that is how the field bootstraps, but it means the benchmarks inherit the biases of their creators. Objects are chosen to make the sensor look good. Tasks are chosen to be tractable on the compute available. Success criteria are chosen to be measurable without human annotation.
The resulting leaderboards are internally consistent and externally uninformative. A policy at the top of a small tactile benchmark rarely outperforms a naive baseline on the same task run on different hardware, and the community has learned to treat the numbers accordingly. What is missing is a benchmark broad enough that its numbers actually predict deployed performance.
Why a benchmark is not the main thing
There is a temptation to say the field's problem is the absence of a good tactile benchmark and to propose one. This is close to right and misses the deeper problem. Benchmarks are downstream of datasets. A serious tactile benchmark presupposes serious tactile training and evaluation data, and that data does not yet exist at the scale required. A leaderboard on top of a shallow dataset produces well-tuned overfitters, not good policies.
The productive order is the opposite. Build the datasets. Publish the datasets. Let the community establish the benchmark protocols on top of the datasets that turn out to be worth training on. This is how vision worked, ImageNet before ImageNet-benchmarks, not the other way around, and it is the order the tactile side of the field is starting to follow.
What the near-term substitute is
In the absence of a serious community benchmark, the near-term substitute is what customers have been doing informally: internal evaluation on the customer's own tasks, using multiple candidate datasets and multiple candidate policies, with a controlled comparison that measures the exact thing the customer cares about. This is less legible than a leaderboard and more informative.
The next piece pulls on the specific reason the absence of a good benchmark is itself a story about the state of the field, not just a technical gap to be filled.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.



