
Vision-Based Tactile Sensors, Explained: GelSight, DIGIT, TacTip in 2026
A field primer on the three vision-tactile families that dominate research fingertips, how they work, what they measure, where they excel, and the practical limits that keep them out of most production datasets.
Vision-based tactile sensors take a very old idea, image a soft surface deforming under load, and modernize it with cheap cameras, elastomers, and neural networks. The result is a class of fingertip sensor that produces images of contact instead of scalar force. Those images turn out to be extraordinarily rich. A single GelSight frame can encode normal force, shear, contact geometry, texture, and incipient slip in a way no strain gauge can.
Three families define the space in 2026. GelSight, born at MIT and now commercialized across several vendors, uses a coated elastomer, controlled internal lighting, and a small camera looking up at the contact patch. DIGIT, from Meta AI, is the open-hardware descendant that made vision-tactile fingertips cheap enough for university labs to buy in bulk. TacTip, from Bristol, uses a different physical principle, tracked pins molded into the elastomer, and produces sparser but more interpretable signals.
All three are excellent research instruments. All three are underused in the datasets that actually train production robot policies. Understanding why requires understanding what they do well and where they break.
How each family works
A GelSight sensor is essentially a tiny photometric-stereo rig. Three or four colored LEDs light a piece of coated silicone from oblique angles. A camera underneath watches the coating deform. Because the lighting geometry is known, the deformation can be inverted into a height map at sub-millimeter resolution, and the height map plus temporal derivatives gives you normal force, shear, and slip. The output is a stream of images, typically at 30 to 90 Hz, and everything downstream is computer vision.
DIGIT preserves the core idea but strips the hardware down to a 3D-printed housing, a Raspberry Pi camera, and an off-the-shelf elastomer. It sacrifices some fidelity for cost and reproducibility. The open design is why DIGIT is now the default vision-tactile fingertip in academic manipulation research.
TacTip takes a different route. Instead of imaging a smooth surface under lighting, it embeds a grid of white pins into the elastomer's underside and tracks the pin displacements as the tip deforms. The signal is discrete and easier to interpret physically, at the cost of spatial resolution. TacTip variants dominate work on tactile servoing and edge-following.
What they measure well
Contact geometry is the standout capability. A GelSight fingertip pressed against a printed circuit board resolves individual solder pads. That is a level of detail no capacitive or piezoelectric fingertip approaches, and it enables tasks, pin insertion, seam following, texture-based part identification, that are effectively impossible with pose-only sensing.
Slip detection is the second standout. Because the sensor images the contact patch over time, incipient slip appears as sub-pixel motion of the surface texture before macroscopic slip occurs. Detecting slip a hundred milliseconds before it happens is exactly the signal a grasp controller needs to increase force preemptively rather than react to a dropped object.
Shear and torque, harder for other sensor classes, fall out naturally from the same image stream. If you can localize contact and track deformation, you can estimate the lateral forces producing that deformation.
Where they hit practical limits
Form factor is the first problem. A GelSight fingertip has to house a camera, LEDs, an elastomer with several millimeters of clearance for imaging, and cabling. The resulting tip is much larger than a human fingertip and much larger than most dexterous robot fingertips. Fitting them onto an anthropomorphic hand costs joint clearance and often knocks out the distal interphalangeal function entirely.
Wear is the second problem. The elastomer coating that makes the sensor work is also the surface doing the mechanical work of the grasp. It abrades. It picks up oils. Its optical properties drift as it ages. Every serious deployment has a coating-replacement schedule, and every replacement requires re-calibration. In a research lab this is a nuisance; in a production capture pipeline it is a throughput killer.
Bandwidth is the third. Camera-based sensors run at video rates. A hundred hertz is aggressive; two hundred is exotic. Contact events, and the human corrections that follow them, contain energy well above that. Vision-tactile sensors filter out the fastest and often most informative parts of the signal.
None of these are fatal in a lab. All of them compound when you need to collect ten thousand hours of skilled human work across dozens of tasks and hundreds of operators.
Where they belong in a full stack
Vision-based tactile is the right choice when you are doing fine-scale research on a small number of tasks, or when you are instrumenting a small number of robots to run policies that were trained on other sensor modalities. It is not the right choice when you are trying to produce large-volume, human-scale, hand-shaped training data.
For that job the field has been quietly converging on a different answer: instrumented gloves that put force and pressure sensing on human fingers themselves, capture at hundreds of hertz, wear like clothing, and scale with an operator network rather than a fabrication line. That is the trade the next two pieces in this series unpack.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.



