
One Million Hours of Human Skill: What a Truly Diverse Corpus Unlocks for Robotics
In August 2026 DYNA-2 proved a human-to-robot scaling law on 1M+ hours of egocentric video. But hours alone are not enough. The technical case for a diverse, force-instrumented million-hour corpus, and why collecting it is Blomega's mission.
One million hours of recorded human skill stopped being hypothetical in August 2026, when Dyna Robotics' DYNA-2 pre-trained a world-action model on more than 1,000,000 hours of egocentric human video and demonstrated the first human-to-robot transfer scaling law: scaling human video predictably improves zero-shot performance on unseen robot hardware. But hours are only one axis. The same scaling research shows diversity beats raw volume, and the modality that decides contact-rich tasks, force, is still missing from the corpus.
What one million hours actually unlocks
DYNA-2, announced on August 10, 2026, is a world-action model pre-trained on over one million hours of human video. It exhibits scaling laws on held-out human data and, critically, a human-to-robot transfer scaling law: more human video predictably improves zero-shot performance on robot hardware the model never saw. Its architecture leans on video co-training, predicting future video states, which the team argues is what enables cross-embodiment generalization.
For a robotics developer the implication is concrete. A large enough human corpus becomes a pretraining substrate you fine-tune with a few hundred robot demonstrations, instead of collecting tens of thousands of teleoperated episodes per task. The corpus is the moat; the robot data is the last mile.
Why diversity beats volume
The second lesson from 2026's scaling research is that raw hours are not the binding constraint. Work on data scaling laws in imitation learning found that policy generalization scales as a power law with data, and that increasing the diversity of environments and objects is far more effective than increasing the number of demonstrations per environment or object.
The mechanism is covariate shift. A policy trained on one kitchen, one set of tools, one operator learns that kitchen, not the task. Generalization comes from coverage of the input distribution the robot will actually meet: new lighting, new clutter, new object instances, new hands. A million hours recorded in ten environments is a smaller effective dataset than a hundred thousand hours recorded across ten thousand.
The diversity axes that matter, and the physics behind each
Countries. Objects, tools, and task conventions are cultural. A cooking or cleaning task in Lagos, Osaka, and Berlin exercises different object geometries, grips, and sequences. Geographic diversity is object and task diversity by proxy.
Tasks. Breadth of skill classes: assembly, folding, wiping, cable handling, food prep. Each has distinct contact dynamics and failure modes, and coverage across them is what makes a policy general rather than task-specific.
Environments. Lighting, surfaces, clutter, and backgrounds drive the vision distribution and the sim-to-real gap. The same task on a matte bench and a reflective counter are different inputs to the policy.
Human gender and body size. This is the axis most datasets ignore and the one that matters most for the hand. Hand span, finger length, and force ranges vary widely across people. A 95th-percentile male hand and a 5th-percentile female hand produce different joint trajectories and different force profiles for the identical task. If your corpus is anthropometrically narrow, your retargeting to a robot hand inherits that bias. Diversity of bodies is what lets normalization map a 45 N grip and a 22 N grip onto the same task-level physics.

The technical spec: what each recorded hour must contain
Video alone, as in DYNA-2, recovers pose and scene but not contact. A corpus built to outlast the video-only era has to carry the physical channels a camera cannot see, synchronized to the frame.
Per recorded hour, at minimum: ten finger-bend channels across two joints on five fingers; fingertip grip force on thumb, index, and middle; a six-axis wrist IMU for trajectory and orientation; all sampled at 100 Hz and hardware-synchronized to a POV or wide video stream over a shared sub-frame time-code. Every channel is timestamped against one clock so force, pose, and pixels align to better than a frame.
Then the layer most groups skip: per-frame calibration and cross-user normalization. Raw sensors drift and vary part to part, so without fleet-scale calibration, hand number 1,847 and hand number 3 do not agree and the data does not merge. Normalization is what turns one person's 45 N grip and another's 22 N grip on the same bolt into a single, embodiment-independent task representation a robot can learn.
Our mission: the diverse, force-instrumented million hours
DYNA-2 proved the ceiling is high. The scaling laws proved diversity is the lever. The force gap proved the video-only corpus is unfinished. Blomega's mission sits exactly at that intersection: collect a million hours of human skill that is diverse by construction, across countries, tasks, environments, and body types, and force-instrumented at the fingertip, not just filmed.
We collect it through a distributed global operator network: real people, in real environments, wearing the capture gear and doing real work, compensated per verified hour through Talika. The gear records the physical channels, the network supplies the diversity, and the calibration layer makes it mergeable. That is how a million hours becomes training data instead of a million hours of video.
Sources and further reading
- DYNA-2: A 1-Million-Hour Scaling Law for World-Action Models, Dyna Robotics (Aug 10, 2026)
- Dyna-2 Proves Scaling Laws for Robotics, Humanoids Daily
- Data Scaling Laws in Imitation Learning for Robotic Manipulation, arXiv 2410.18647 (ICLR 2025)
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data, arXiv 2602.16710
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.



