
The Schema War in Robot Learning: RLDS vs LeRobot, and Who Actually Wins
Robot-learning data is consolidating around two formats: Google DeepMind's RLDS and Hugging Face's LeRobotDataset. LeRobot has the momentum, but the winning schema still under-specifies force.
As of 2026, robot-learning data is consolidating around two formats: Google DeepMind's RLDS, the TensorFlow-based backbone of Open X-Embodiment, and Hugging Face's LeRobotDataset, a PyTorch-native format carrying 16,000-plus datasets from 2,200-plus contributors. LeRobot has the community momentum. But both under-specify the one signal that decides contact-rich tasks: force.
What is the robot-data schema war?
Every robot-learning dataset needs a container: a standard way to store synchronized camera feeds, joint states, actions, and task labels so any team can load and mix them. The schema war is the competition over which of those containers becomes the default the whole field trains on. Own the format and you sit under everyone's pipeline.
RLDS vs LeRobot: who is ahead?
RLDS, from Google DeepMind, is built on TensorFlow Datasets and Apache Arrow and is the format all of Open X-Embodiment is converted to, making it the pre-training backbone for RT-1-X, RT-2-X, and Octo. LeRobotDataset, from Hugging Face, is PyTorch-native and lower-friction, and it has taken the open-source lead with more than 16,000 datasets from over 2,200 contributors as of September 2025.
By format: RLDS is backed by Google DeepMind on TensorFlow plus Apache Arrow, powering Open X-Embodiment, RT-X, and Octo. LeRobotDataset is backed by Hugging Face, is PyTorch-native, and carries 16,000-plus datasets from 2,200-plus contributors as of September 2025. InternData-A1, from InternRobotics, competes on synthetic scale.
Is the fragmentation era ending?
One ecosystem analysis argues the convergence of Open X-Embodiment for TensorFlow-scale unification, LeRobot for PyTorch-native accessibility, and InternData-A1 for synthetic scale signals the end of the fragmentation era in robot-learning data. Cross-embodiment portability, not a single winner, is the direction of travel.
What the winning schema still misses: force
Whichever container wins, it will still be vision-and-proprioception first. Recent surveys note that compared with vision, tactile sensing remains far less standardized and far less widely used, especially for dexterous and contact-rich manipulation. The fields for fingertip force, grip, and slip are barely specified in the formats now competing to win.
Our analysis
The schema war will be won on vision and proprioception, but the format that wins will still under-specify force. The real opening is not a new container, it is the force and tactile fields the current standards barely define. Blomega Lab captures fingertip force and bend synced to POV video and exports it into whatever format wins, LeRobot or RLDS, rather than betting on the container. Own the modality, not the war.
Sources and further reading
- LeRobot: An Open-Source Library for End-to-End Robot Learning, arXiv 2602.22818
- Open X-Embodiment / RT-X and the RLDS format, Google DeepMind
- Open Robotic Data at Scale: Ecosystem Formation and Implications, Codatta (Dec 2025)
- Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning, arXiv 2608.07558
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.


