Notes from the data layer of robotics.
Education, research, and field notes. Written for engineers building robots and the operators teaching them.

The Egocentric Data Wave: 2026's Biggest Robot-Learning Datasets and the Modality They Miss
In 2026 egocentric human video hit scaling laws for robot manipulation, from EgoScale's 20,854 hours to EgoVerse's 1,362. But the one signal it cannot capture is force at the hand.

The Schema War in Robot Learning: RLDS vs LeRobot, and Who Actually Wins
Robot-learning data is consolidating around two formats: Google DeepMind's RLDS and Hugging Face's LeRobotDataset. LeRobot has the momentum, but the winning schema still under-specifies force.

One Million Hours of Human Skill: What a Truly Diverse Corpus Unlocks for Robotics
In August 2026 DYNA-2 proved a human-to-robot scaling law on 1M+ hours of egocentric video. But hours alone are not enough. The technical case for a diverse, force-instrumented million-hour corpus, and why collecting it is Blomega's mission.

Open Developer Ecosystems for Robotics: A Technical and Commercial Field Guide (2026)
A deep reference on what an open developer ecosystem for robotics actually is in 2026: the middleware, simulation, data, and model layers; the business models and robot app stores; the open-versus-closed data-moat tension; and what has to be true for an iOS or Android moment in physical robots.

The Robot Data Pipeline: A Unified Sensor-to-Cloud Recording and Upload Strategy for Navigation and Perception
A deep technical reference on moving robot data from sensor to cloud: what to record and why you cannot upload it all, MCAP and triggered recording on the edge, time-synchronizing navigation and perception, curating the critical one percent, and the offload, storage, and flywheel that follow.

The Robot Manipulation Dataset Registry: 19 Datasets, and Only 2 Have Force
A cleaned, normalized, machine-readable registry of open robot manipulation datasets. The scattered landscape refined into one schema, and one column that tells the story: only 2 of 19 datasets carry force at the hand.

How Much Force Does It Take? A Registry of the Newtons Behind Everyday Tasks
A cleaned, sourced registry of the force and torque humans apply to everyday manipulation tasks, from 0.45 N to press a key to ~5 Nm to open a jar. The one axis a camera cannot see, refined from the ergonomics literature into agent-ready data.

The Tactile and Force Sensor Registry: 24 Ways to Give a Robot Touch
A cleaned, sourced registry of the tactile and force/torque sensors robots use, from open-source GelSight and ReSkin to industrial ATI and Bota. One normalized schema, and a split that tells you where the field is going.

Who Builds Robot Manipulation: A Registry of 39 Labs and Companies
A cleaned, sourced registry of the academic labs and companies building robot manipulation and the data behind it, from Stanford IRIS and Berkeley RAIL to Physical Intelligence, Figure, and AgiBot. One schema for the whole field.

The Robot Foundation Model Registry: 18 VLAs and Policies, and Which Are Open
A cleaned, sourced registry of the vision-language-action models and manipulation policies that run on robots, from 27M-parameter Octo to 55B RT-2-X. Two-thirds ship open weights; the biggest company results do not.

End-to-End Robot Training: How Robots Learn Skills in 2026 (The Ground-Truth Guide)
End-to-end robot training replaces the hand-coded perception-planning-control stack with one learned model that maps sensors straight to actions. This is the complete, sourced guide: the pipeline, the three learning methods, all 18 foundation models compared, where the data comes from, sim-to-real, and the contact-force gap holding it back.

Where BLO LAB Sits in the Coordinated Strategy
A summary of the position: wearable-first capture, a global operator network via Talika, a delivery pipeline built for foundation-model training, and honest complementarity with the teleop-warehouse layer.

The Geography of the Operator Network
Where the operators live is not a footnote. It determines what skills are cheap to capture, what compliance regime applies, and what languages your annotation pipeline has to speak.

The Hundred-Million-Hour Question
A universal manipulation policy will eventually need something like a hundred million hours of demonstration data. Nobody in the market can currently produce that. The path to it is the real strategy question.

Data Services vs Data Software: Two Different Businesses
A company that sells operator-hours of demonstration is a services business. A company that sells a capture platform is a software business. Both are legitimate. They are not the same company.

The Pick-and-Shovels Thesis for Physical AI
In a gold rush, the reliable business is not gold. It is the picks and shovels. In the current robotics cycle, capture hardware and operator networks are the shovels, and they are being bought.

Provenance and Audit for Manipulation Datasets
As foundation-model labs get bigger and more regulated, the audit trail on their training data becomes a first-class concern. Data vendors that cannot produce one will be filtered out.

The Emerging Standard for Manipulation-Data Delivery
There is no formal standard yet for how a manipulation dataset should be shipped. There is, informally, a converging one, and vendors that ignore it pay the price in integration weeks.

What Frontier Labs Actually Buy When They Buy Data
Nobody buys 'data.' They buy specifications, delivery formats, QA guarantees, and the right of first refusal on the next capture cohort. Understanding the actual purchase order is where vendors either win or lose the account.

OpenAI's Robotics Relaunch and What It Means for the Data Layer
The largest AI lab in the world reopening its robotics program tells us where the demand curve is going. What it does not tell us is where the supply is going to come from.

Should a Foundation-Model Lab Buy Data or Build It?
Every lab building a physical-AI foundation model faces the same make-or-buy question on data. The right answer depends on the shape of the model, not just the size of the budget.

Wireless, Battery-Powered, All-Day: The Full-Body Capture Spec
A suit that only works tethered, in a studio, for thirty minutes at a time captures a fraction of the work that matters. The full-body capture spec is what unlocks the rest.

Retargeting Is the Strategic Layer of Wearable Capture
The mapping from human motion to robot action is not a lossy compromise. Done right, it is the layer that lets one dataset feed every embodiment a lab ever ships.

Contributing a Force-Rich Reference Dataset to the Field
The BLO LAB reference dataset is a curated slice of our production capture, released with force, pose, and video aligned to the standards a serious tactile benchmark will eventually require. Here is what it contains, why we ship it, and what we hope the field does with it.

Why Force and Tactile Signal Belongs at the Fingertip
A follower arm can approximate force from joint torque. A fingertip sensor measures it directly. For contact-rich policies, the gap between the two is the gap between a policy that works and one that plateaus.

In-Situ Capture: Data Where the Work Actually Happens
The kitchen, the workshop, the lab bench, these are the places skilled work happens. A capture stack that has to be shipped to them beats one that expects them to come to it.

Why Wearable Capture Scales on a Different Curve
The bottleneck in wearable data collection is skilled operators, not arms, cells, or floor space. That shifts the entire growth curve of the business.

The Embodiment Transfer Problem in Teleoperated Data
Data captured on one robot does not automatically train a policy for another. The gap between capture embodiment and deployment embodiment is the hidden tax on every teleoperation dataset.

Cost Per Demonstration Is the Only Metric That Matters
Total dataset cost is the wrong denominator. The metric that decides whether your data program compounds is the fully-loaded cost of one usable demonstration hour.

Robot-in-the-Loop vs Human-in-the-Loop Capture
Every manipulation dataset is captured in one of two modes. The tradeoffs between them govern what you can teach a policy, how fast, and at what price.

The GELLO Legacy: What Comes After the $300 Puppet Arm
The GELLO teleoperation rig lowered the cost of imitation-learning data collection by an order of magnitude. Its own authors are now building the company that has to answer what comes next.

There Is No Serious Tactile Benchmark. That Is the Story.
The absence of a broadly-adopted tactile benchmark is not an oversight. It reflects how thoroughly the field is still operating on vision-first assumptions. Naming that gap is more useful than trying to fill it prematurely.

The Teleoperation Ceiling: Why Warehouses of Robot Arms Cap Out
Robot-in-the-loop capture, an operator flying a leader arm that drives a follower arm, is the fastest way to start collecting manipulation data. It is also the mode that hits a wall first.

What a Serious Tactile Benchmark Would Look Like
The field has vision benchmarks, manipulation benchmarks, and mobility benchmarks. It does not have a serious tactile benchmark. A design sketch for what one would need to include, and why building it is more useful than another leaderboard.

The Operator Network Is Infrastructure
Foundation-model labs talk about compute, data, and hardware. The fourth pillar, the humans who wear the capture rig and do the work, is the one nobody has industrialized yet.

Aligned Multi-Stream Capture: The Wearable Advantage
The single biggest advantage of a wearable capture rig over a teleop-warehouse or lab setup is what happens between modalities. Sub-millisecond alignment across force, pose, video, and biomechanics is not an add-on, it is what makes genuinely multimodal training possible.

Simulation Cannot Fake a Contact Event
High-fidelity simulators have become impressive enough that some teams argue real-world capture is optional. For dexterous manipulation, the argument breaks the moment two surfaces touch.

Most 'Multimodal' Robot Models Are Still Vision With Extras
The multimodal framing hides how thoroughly vision dominates modern robot policies. A frank look at the input mix in the most-cited systems, why the imbalance persists, and what a genuinely modality-balanced policy would require.

The Universal Robot Brain Has a Universal Data Problem
A single foundation model that drives any robot on any task is the most ambitious bet in physical AI. Its ceiling is not compute, it is the breadth and fidelity of the demonstrations you can feed it.

Vision + Force + Proprioception: The State of Multimodal Robot Learning
Multimodal is one of the most-used words in robot learning and one of the least examined. A working survey of what modalities are actually being fused in 2026, what the fusion architectures look like, and where the honest advances are.

Capturing Human Compliance as a Training Signal
Human hands are compliant instruments. Capturing what they do, not just where they go, turns a skilled operator into a source of the exact training signal that stiffness-blind robot policies lack. Here is how the capture works in practice.

Why Every Serious Robotics Lab Is Building Its Own Hand
Genesis AI shipped a proprietary dexterous hand alongside its foundation model. So did most of its peers. The reason is not vertical integration for its own sake, it is that the hand defines what the data has to look like.

Compliance Policies Are Undertrained Because Compliance Data Is Rare
The robotics community has known for years that compliant behavior is essential and that most learned policies do not exhibit it. The reason is upstream of any model architecture, the training data does not contain the signal.

Long-Horizon, Contact-Rich: Why Cooking Broke Robotics
Making a smoothie, harnessing a wire, or solving a Rubik's cube are the tasks foundation-model labs now benchmark on. They share the two properties that classical robotics could never handle at the same time.

A Roadmap to General-Purpose Household Robots
Where the field is, where it is going, and what has to be true for household robots to ship.

Impedance and Admittance Control, Without the Math
Compliance is the robotics word for behaving softly. A plain-language walkthrough of impedance and admittance control, why they matter for contact-rich tasks, and how they interact with modern learned policies.

The Economics of Robotics Data: Cost per Demonstration
Understand the unit economics before you scale.

Building a Slip-Labeled Dataset From Real Human Work
How BLO LAB captures, labels, and delivers slip and near-slip events at the scale foundation-model training requires. A concrete look at the collection protocol, the physics-based labeling, and what the resulting dataset lets a policy learn.

Evaluating Robot Policies: Metrics That Actually Matter
Success rate is not enough. Here is what to measure instead.

Slip Is the Failure Mode Nobody Trains For
Robot grasping benchmarks report success rates. Real deployments care about failure modes. Slip is the dominant failure mode in field data and the least represented in training data, a mismatch that quietly caps every policy shipped today.

Building a Multi-Modal Dataset: Video, IMU, Force, Audio
Combining streams is where the real work is. A field guide.

Slip Detection in Robot Grasping: A Field Primer
Slip is the physical event that decides whether a robot keeps hold of what it picked up. A working guide to how the field detects it, why traditional methods break, and what a modern slip-aware policy looks like.

From Contributor to Paycheck: How Talika Pays for Human Skill
The economics of the contributor side. How sessions turn into real income.

The 65% Problem: Why Fine Motor Skills Still Belong to Humans
Across advanced manufacturing, roughly two-thirds of remaining human labor is there for one reason, fine motor skill. Understanding that number is understanding where robotics goes next.

Force-Rate Capture on the BLO LAB Glove: The Signal Chain
A walk through the BLO LAB glove's force pipeline from sensor to serialized frame. What we sample, at what rate, how we calibrate across operators, and how the resulting stream survives the trip into a foundation-model training loader.

Whole-Body Manipulation and Why Suits Beat Cameras
Real tasks recruit the whole body. Capture that or leave capability on the table.

The 90% of Manipulation Datasets That Skip Force Entirely
A frank audit of the public and semi-public manipulation datasets that dominate foundation-model training. Force appears in a small minority, is calibrated in a smaller minority, and is present at usable rates in a smaller minority still. The gap is a strategic opening.

Folding, Peeling, Kneading: The Long Tail of Household Tasks
General-purpose robots die on the long tail. Coverage is the only way through.

Capturing the Craftsman: Turning Skilled Workers Into Training Data for Physical AI
The most valuable training data in robotics is not on the internet. It is inside the hands of the workers who already do the task. Here is how a modern capture program brings that skill into a model.

Why Fingertip Force Is a Better Signal Than Fingertip Pose
Pose tells you where the finger was. Force tells you what the finger was doing. For contact-rich tasks, which is most of them, the second is decisive and the first is derivative. A working primer on why the field is finally taking force seriously.

Kitchen Robotics: Why Cooking Is the Ultimate Benchmark
Kitchens are chaotic, deformable, and unforgiving. That is exactly why they are the right frontier.

Why We Chose Force-Instrumented Gloves Over Vision-Tactile Fingertips
The BLO LAB glove was designed against a specific trade-off: fingertip vision-tactile sensors are richer per contact, but instrumented gloves scale to the volume, embodiment, and task breadth a foundation-model dataset actually requires. Here is the reasoning, in full.

Data Licensing for Robotics: What Buyers and Contributors Should Know
Licensing terms decide who can train, what they can ship, and how contributors get compensated.

Vision-Tactile Sensors Are Beautiful. They're Almost Never in Real Datasets.
A survey of the largest public manipulation datasets shows vision-tactile signals in a vanishing fraction of demonstrations. The gap between what these sensors can do and what production data pipelines actually contain is where the field is quietly losing capability.

4D+ Motion Capture: Why Time-Aware Hand Data Beats Pose Snapshots
Traditional mocap gives you pose over time. 4D+ mocap gives you pose, force, contact, and micro-timing, the four axes a dexterous policy actually needs.

The Case for Egocentric Video in Foundation Models for Robotics
First-person video captures intent and attention in a way third-person cameras never will.

Vision-Based Tactile Sensors, Explained: GelSight, DIGIT, TacTip in 2026
A field primer on the three vision-tactile families that dominate research fingertips, how they work, what they measure, where they excel, and the practical limits that keep them out of most production datasets.

Understanding Degrees of Freedom in Robotic Hands
DoF is thrown around casually. It hides more nuance than people admit. A short primer.

Dexterity-First Foundation Models: What a Robotics Foundation Model Actually Requires
The field is converging on a new class of model, the Robotics Foundation Model, trained not on text but on trajectories. Here is what makes a dexterity-first RFM different from a generalist VLA.

Anatomy of a Motion-Capture Suit for Robotics (Roadmap)
What a robotics-first mocap suit should look like, and why we're shipping the glove first. A design brief, not a datasheet.

Inside the Glove: The Sensors Behind GX-1
A hardware tour of the honest v1 capture glove: 10 bend channels, fingertip force on three fingers, and a 6-axis wrist IMU, all 100 Hz hardware-synced.

Dexterity Is the Last Mile of Industrial Automation
Robots have automated the heavy, repetitive middle of manufacturing. What is left is the human hand, and closing that gap is now the single largest opportunity in physical AI.

Sim-to-Real Transfer: Where It Breaks and How Real Data Fixes It
Simulation is fast and cheap. It is also wrong. Here is where the reality gap opens and what to do about it.

How to Build a Dexterous Manipulation Dataset from Scratch
A practical checklist for teams standing up their first manipulation dataset, drawn from real deployments.

Diffusion Policies Explained for Robotics Engineers
Why diffusion, borrowed from image generation, quietly took over robot policy learning.

Imitation Learning 101: From Human Demonstration to Robot Policy
A walkthrough of how a recorded human demonstration becomes a running robot policy, without the math jargon.

Why Force Data Is the Missing Ingredient in Robot Manipulation
Robots that only see fail at soft objects, deformables, and anything requiring subtle grip. Force data solves that.

Motion Capture vs. Vision-Only Learning: A Practical Guide
Cameras are cheap, but they miss the physics. Here is when to reach for motion capture and when video is enough.

What Is Teleoperation Data and Why Robotics Needs It
A plain-English introduction to teleoperation, how it produces training data, and why every serious robotics team is investing in it.
