
Capturing the Craftsman: Turning Skilled Workers Into Training Data for Physical AI
The most valuable training data in robotics is not on the internet. It is inside the hands of the workers who already do the task. Here is how a modern capture program brings that skill into a model.
There is a specific kind of knowledge that lives only in hands. A line cook can tell whether dough is ready by the resistance under the heel of the palm. A cable technician can seat a connector by the click their fingers hear before their ears do. An upholsterer knows the exact angle to pull a staple so it does not bend. None of this is written down. Very little of it is teachable in words. All of it is exactly what a physical-AI system needs to learn.
For decades this kind of skill was invisible to software. The workers who had it did the work; the work happened; nothing about it entered the digital record. The rise of dexterity-first robotics has changed that calculus. Suddenly the trained hand of an experienced worker is a source of training signal that no simulator, no scraped video, and no synthetic pipeline can substitute for.
The programs that turn that skill into data are new, and worth describing carefully because they are becoming a real category of work.
Recruiting for the hand, not the resume
A capture program looks nothing like a data-labeling program. Labeling recruits for volume and consistency; capture recruits for skill. The right contributor for a folding dataset is not the cheapest available operator; it is a hotel housekeeper who has folded ten thousand sheets. The right contributor for a wire-harness dataset is a technician who has actually built harnesses.
This changes the economics of the operator network. You cannot scale it purely by adding seats. You have to source people who bring the underlying skill, and you have to pay them for it. Rates that would be absurd for click-work, twenty, thirty, sometimes fifty dollars an hour, are entirely reasonable for a skilled operator producing high-value trajectories a foundation-model team will use for years.
Instrumenting the task without changing it
The second design principle is that the capture rig has to disappear. The moment an operator has to think about the sensors, the trajectory stops looking like real work and starts looking like a demo. This is why wearable capture, gloves, straps, headband cameras, dominates the space. Optical mocap requires a stage; wearable capture goes to where the work already happens.
The bar for wearability is high. A glove that changes how the operator grips loses the very skill you were trying to record. A suit that limits shoulder range narrows the trajectory distribution. Every gram, every constraint, every wire that gets in the way is a subtle distortion of the training signal.
This is why the frontier hardware in the space looks the way it does: as light as possible, as unobtrusive as possible, with sensing distributed so the operator forgets it is there within the first ten minutes.
Turning raw capture into training-ready data
A capture session produces a bundle of channels, pose, force, contact, video, audio, sometimes EMG, that has to be cleaned, synchronized, segmented, and labeled before any model consumes it. This is where most of the actual work in a capture program sits, and it is where most amateur efforts fail.
Segmentation matters most. A one-hour session of folding towels is not one datapoint; it is often a hundred and fifty distinct trajectories, each with a clear start and end. Getting those boundaries right, consistently, across thousands of hours and hundreds of operators, is the job of a real data pipeline. Automating what can be automated, and having skilled human reviewers on what cannot, is the difference between a dataset that trains a policy and a directory of noisy files.
Paying contributors like the specialists they are
The last piece, the one that decides whether a capture program is sustainable, is compensation. Contributors who produce high-quality trajectories deserve to be paid for the skill they brought, not just the time they spent. Payment models that pay per usable trajectory, or per hour of clean capture after review, align the operator's incentives with the buyer's. Payment models that pay flat hourly regardless of quality do not, and they produce datasets that reflect it.
This is the economic architecture behind Talika. Skilled contributors wear the gear, do the work they already know how to do, and get paid per hour of usable capture. The dataset that results is not a scrape of the internet or a rendering from a simulator. It is a direct recording of skilled human labor, licensed to the teams building the models that will one day do that labor at scale.
Building or training robots?
We license manipulation datasets and run custom capture programs. Get in touch to see what fits.



