



































































































When we say diverse, we mean diverse.
Every tile behind this is real capture. The same task in a different kitchen, a different country, different hands, different light. Breadth is the data priority robots need, and we cover every axis of it.
A global operator network
Every region brings its own tasks, objects, and tools. Hover or tap a location.


The same activity, done the local way. Same capture standard everywhere: head-mounted POV, hands in frame, real practitioners, no PII.
Task distribution across the network
Not one big pile. Structured families of work, each done differently per country. Pick a family to see the local variants and light it up on the map.
One demonstration =
Click any tile to cycle that axis, or shuffle the whole configuration. The preview follows the task.
Real tasks, real hands
A sample of live capture. Every clip is one point in the space above.
Environment coverage
89 real-world verticals across 16 categories, prioritized per project. Additional verticals onboarded on request.
Food & Kitchen
7Food Production & Processing
4Vehicle & Mobility Services
7Building & Trades
6Maker & Craft Workshops
5Industrial Manufacturing
7Consumer Product Manufacturing
2Print & Visual Production
2Repair & Maintenance
6Retail & Storefront
15Apparel & Fashion
6Hospitality & Lodging
3Sports & Leisure
2Industrial & Warehouse
8Household (non-cooking)
8Dental Laboratory
1The capture standard
Diversity only counts if every clip meets the same bar. Here is the standard behind each one.
- Mono
1920x1080, 30 fps hard CFR, H.264 High, 100 Hz IMU sidecar, per-frame camera intrinsics, AAC 44.1 kHz audio
- Stereo
4000x1200 (1920x1200 per eye), 30 fps, 4 sensor sidecars, 160° diagonal FOV (H 134°, V 72°), configurable 100° to 200°
- Head-mounted first-person view, wide-angle
- Camera at forehead level, angled about 45° downward
- Both hands, both feet, and the workspace visible in frame
- Contributor standing, real task at natural pace
- Continuous execution, hands out of frame under 10% of grab/release time
- Per-clip metadata delivered as a CSV/JSON index alongside footage
- Reviewed within one day of submission, rejected clips re-captured
- Real practitioners doing real work, no staged or scripted activity
- No minors, no PII, authorized locations only
- Consented and rights-cleared, full rights assignment to the buyer
Every axis we vary
18 dimensions of variation. Additional values available on request, new verticals onboarded per project.

Tasks & actions
46The verb itself. Manipulation is thousands of distinct skills, not one.















































Objects & materials
32Rigid, deformable, articulated, granular, liquid. Each behaves under contact differently.

































Environments & verticals
89Where the task actually happens. Real worktops, real clutter, real constraints.






















































































Lighting & visual conditions
17Vision policies break under lighting shift. Coverage here is coverage of the real world.

Handedness
3Left and right change the whole demonstration geometry.

Hand & body size
8Anthropometry shifts viewpoint, reach, and grasp. Models must not overfit one body.

Skin tone
7Full representation across skin tones so hand tracking works for everyone.

Operator experience
8Master to novice. Expert paths and beginner mistakes are both signal.

Age range
6Adults across the full range (no minors). Age changes speed, grip, and technique.

Gender
3Balanced representation across the operator network.

Country & region
30A global operator network. The same task looks different across the world.

Cultural technique variant
8The same goal, done the local way. Regional methods, tools, and conventions.

Pace & timing
8Deliberate to production-speed. Real tempo, including pauses and interruptions.

Task horizon
5A 2-second action to a session-length workflow. Long-horizon structure is rare and valuable.

Contact & precision regime
8Delicate to forceful, coarse to millimeter-precise. The physics of the grip.

Outcome coverage
8Not just clean successes. Corrections, near-misses, and recoveries are the data models lack most.

Capture modality
8How the demonstration is recorded. Layered from POV video up to force.

Action labels & sidecars
12What makes it trainable, not just footage. The layer raw egocentric video lacks.
Image credits, 209 photographs, Creative Commons via Openverse
Diversity you can audit, demonstrations you can train on.
Raw egocentric video is not enough. It carries no metric 3D hand trajectory, no action labels, no clean task structure. We deliver the missing layer: diverse, real human demonstrations across every axis above, each one action-labeled and specified, so your policies learn the breadth of the world instead of one corner of it.
