All articles
Research·7 min read·August 14, 2026

The Egocentric Data Wave: 2026's Biggest Robot-Learning Datasets and the Modality They Miss

In 2026 egocentric human video hit scaling laws for robot manipulation, from EgoScale's 20,854 hours to EgoVerse's 1,362. But the one signal it cannot capture is force at the hand.

As of August 2026, egocentric human video has become the dominant way to scale robot manipulation data. EgoScale trained a vision-language-action model on 20,854 hours of action-labeled egocentric video and found a log-linear scaling law; EgoVerse released 1,362 hours across 1,965 tasks. The one signal none of them capture is force at the hand.

Close-up of the BLO LAB GX-1 sensing glove capturing fingertip force during a manipulation task
The GX-1 records fingertip force and finger bend, the contact signal egocentric video alone cannot see.

What is EgoScale and why does it matter?

EgoScale, from the UT Austin Robot Perception and Learning Lab, trained a vision-language-action model on more than 20,854 hours of action-labeled egocentric human video, over 20x larger than prior efforts, and uncovered a log-linear scaling law between human-data scale and validation loss, which correlates with downstream real-robot performance.

How big is the new egocentric data?

2026 produced a wave of large, collaborative egocentric datasets. EgoVerse alone spans 1,362 hours across 1,965 tasks, 240 scenes, and 2,087 demonstrators, recovering 21 hand keypoints and 6-DoF head pose via visual-inertial SLAM.

By dataset: EgoScale covers 20,854 hours for VLA pretraining and a scaling law (arXiv 2602.16710); EgoVerse covers 1,362 hours across 2,087 demonstrators of in-the-wild manipulation (arXiv 2604.07607); EgoHumanoid pairs Unitree G1 teleoperation with egocentric demos for whole-body loco-manipulation (RSS 2026).

Why egocentric video, and why now?

Human egocentric video is far cheaper to collect than robot teleoperation, covers a vastly larger range of tasks and environments, and, when aligned correctly, transfers to robot policy performance. That economics is what pushed egocentric data into the mainstream of robot learning in 2025 and 2026.

What can egocentric video not capture?

Egocentric video recovers where the hand is, not how hard it presses. Contact force, grip, and slip are a separate, scarcer signal, an active but distinct research line in tactile foundation policies. A camera on the head cannot see the newtons at the fingertip.

Our analysis

The egocentric wave proves that scaling human data works for pose and vision. But the scaling law that has not been run yet is the one on force. The modality a camera cannot capture is exactly what the Blomega Lab GX-1 records: fingertip force and bend, hardware-synced to the same POV video these datasets already use. As egocentric pretraining saturates, force-conditioned data is the next axis of improvement.

Sources and further reading

Work with us

Building or training robots?

We license manipulation datasets and run custom capture programs. Get in touch to see what fits.