
There Is No Serious Tactile Benchmark. That Is the Story.
The absence of a broadly-adopted tactile benchmark is not an oversight. It reflects how thoroughly the field is still operating on vision-first assumptions. Naming that gap is more useful than trying to fill it prematurely.
In vision, the community has ImageNet, then a decade of successor benchmarks, and now a stack of foundation-model evaluations that everyone at least argues about. In language, the equivalents are more contested but they exist. In robot manipulation, there are enough benchmarks that comparing across them is a small research topic in itself. In tactile learning, there is no equivalent. It is worth asking why, because the answer is not that no one has tried.
What has been tried
Several labs have proposed tactile-focused benchmarks over the last five years. Most were tied to a specific fingertip design or a specific task set. Most produced leaderboards with a handful of entries and then went quiet. None became the coordination point the field organizes around. This is not a story of failed effort; it is a story about what a benchmark needs to become a field-wide standard, and what tactile lacks.
Why a tactile benchmark is structurally harder to establish
Vision benchmarks succeeded because vision data is cheap, portable, and easy to standardize. An image is an image regardless of which camera captured it, roughly speaking, and the discrepancies are small enough that a benchmark can absorb them.
Tactile data is none of these things. A GelSight frame, a DIGIT frame, a TacTip frame, a piezoresistive glove frame, and a capacitive fingertip frame all measure different physical properties, at different sampling rates, at different spatial resolutions, in different units. Standardizing across them requires either an abstraction layer that discards information or a benchmark protocol that runs on multiple sensor types simultaneously. Both are hard, and neither has happened.
The second reason is that tactile evaluation requires real hardware in a way vision evaluation does not. You can benchmark a vision model on a static image dataset. You cannot benchmark a tactile policy on a static tactile dataset in any deep sense, because tactile behavior is defined by the closed-loop interaction between action and sensed contact. A benchmark that runs on offline data measures something different from what deployment measures, and the community has been unwilling to let the difference slide.
What the absence signals
The absence signals that tactile is still in a phase where the shape of the field is being negotiated. Which sensor modalities matter. Which tasks are the right ones. What the evaluation protocol should look like. These questions have not settled, and it is not obvious they should be forced to settle before more of the underlying work has been done.
It also signals that vision-first assumptions have persisted longer than they should have. If tactile were treated as an equal partner in manipulation research, the field would have organized around a tactile benchmark by now, in the same way it organized around cross-embodiment benchmarks once the community decided cross-embodiment mattered. The absence is diagnostic.
Why the story is going to change
It is going to change because deployment pressure is going to force it. The customers paying real money for robot policies are not evaluating those policies on vision benchmarks. They are evaluating them on their own tasks, and their tasks are increasingly the contact-rich tasks that vision-only policies do poorly on. The internal evaluations at foundation-model teams and at customers have started to look tactile-aware even where the public benchmarks have not caught up.
The community benchmarks will follow. They usually do. When they arrive they will look different from the vision benchmarks that preceded them because the underlying problem is different, and the labs that shaped the datasets in the intervening period will shape the benchmarks that eventually sit on top of them. This is one of the reasons a serious dataset supplier participates in the pre-benchmark period rather than waiting for it to end.
What we are doing in the meantime
We publish reference datasets, ship them with documented protocols, and support customers in running their own controlled evaluations on the tasks they care about. When a broadly-adopted tactile benchmark eventually appears, the datasets and protocols will be in a position to inform it. Until then, the substitute for a leaderboard is a clean internal comparison, and the substitute for a benchmark is a well-scoped customer evaluation.
The next piece describes the specific reference dataset we contribute back to the community as a step toward the benchmark that does not yet exist.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.



