Demonstration videos make it look as though robots learn a task by being shown it once. They do not. What changed over the past few years is less cinematic and more interesting: large shared datasets collected across many different machines, policies borrowed from image generation research, and simulators used in a way that tolerates being wrong. This article walks through the published work behind those shifts, and is clear about which parts remain genuinely hard.
Updated October 2026. Figures below are those reported by the papers and developers cited, under their own evaluation conditions.

From programming to imitation
Classical industrial automation does not learn anything, which is the baseline against which modern claims that robots learn should be read. A path is specified, the environment is controlled, and repeatability is the whole point. That works superbly for a car body on a line and fails as soon as the object is in a slightly different place or made of something that deforms.
Imitation learning replaced the explicit path with examples. A human teleoperates the arm through a task many times, and a neural network learns to map what the cameras see to what the joints should do next. The quality of the result depends on the quality and variety of those demonstrations, which is why data collection, not algorithm design, is where most of the effort in the field now goes.
The borrowed idea that made policies much better
One of the clearest step changes came from treating action prediction as a generative problem. Diffusion Policy represents a robot’s visuomotor policy as a conditional denoising diffusion process, the same family of methods behind image generators. Evaluated across 12 tasks from 4 robot manipulation benchmarks, the paper reports an average improvement of 46.9 percent over prior state of the art.
The reason this matters conceptually is that a diffusion process handles ambiguity well. There is rarely one correct way to reach for a mug, and a method that can represent several plausible continuations behaves better than one forced to average them into a single bad compromise.
How robots learn from each other’s data
The second shift was pooling data across machines. The Open X-Embodiment project assembled a dataset covering 22 different robots collected through a collaboration between 21 institutions, spanning 527 categories of behaviour across 160,266 tasks. The models trained on it, called RT-X, are reported to exhibit positive transfer and to improve the capabilities of multiple robots by drawing on experience from other platforms.
That is a significant claim, because the long-standing assumption was that every robot needed its own data. More recent work pushes the same idea further. The pi-0 model is described as a flow matching architecture built on a pre-trained vision-language model in order to inherit internet-scale semantic knowledge, trained on data from single-arm, dual-arm and mobile manipulators, and evaluated on tasks including laundry folding, table cleaning and assembling boxes.
Google DeepMind’s Gemini Robotics work frames the same direction around three properties: generality, interactivity and dexterity. The company reports that the model more than doubles performance on a comprehensive generalisation benchmark compared with other vision-language-action models it considers state of the art. Treat benchmark figures as what they are, measurements under the authors’ own conditions, and the direction as the real signal. The pattern mirrors what happened in language work, which our overview of how AI is changing everyday work traces on the software side.
Simulation, and why it is deliberately wrong
Simulators are cheap and fast, and they are never accurate about the things that matter most in manipulation: friction, contact and deformation. The classic answer is domain randomization, described in the original paper as “a simple technique for training models on simulated images that transfer to real images by randomizing rendering in the simulator”. The paper reports a real-world object detector trained entirely on randomised synthetic images achieving localisation accuracy to 1.5 cm, with no pre-training on real photographs.
The logic is counterintuitive and worth stating plainly. If the simulation varies wildly in appearance, the real world becomes just another variation rather than an unfamiliar case. Randomising textures and lighting is now routine; randomising physical properties such as mass and friction is harder and is where much of the remaining gap sits.
6 simple ideas that explain how robots learn
- Demonstrations, not instructions. A policy is trained from recorded human teleoperation, so the data defines the behaviour far more than the code does.
- Shared data across different bodies. Open X-Embodiment pooled data from 22 robots and 21 institutions and reported positive transfer between platforms.
- Generative action prediction. Diffusion Policy models actions as a denoising process and reported a 46.9 percent average improvement across 12 tasks and 4 benchmarks.
- Language and vision as a starting point. Models such as pi-0 build on pre-trained vision-language models to inherit general semantic knowledge before any robot data is added.
- Randomised simulation. Training on deliberately varied synthetic scenes makes reality one more variation; the original work reported 1.5 cm localisation accuracy with no real training images.
- Measurement as the bottleneck. NIST develops performance metrics and test methods for robotic grasping and assembly precisely because comparing manipulation results across systems is hard.
What is still genuinely difficult
- Contact-rich manipulation. Anything involving force, friction or a deformable object remains far harder than navigation or pick-and-place with rigid items.
- Recovery from mistakes. Demonstrations mostly show success, so policies see few examples of getting out of trouble.
- Comparable evaluation. A number from one laboratory rarely transfers to another, which is the problem NIST’s test methods address.
- Reliability at length. A 95 percent per-step success rate is poor across a twenty-step task, and household work is full of long tasks.
- Honest demonstrations. Before believing a video, ask whether it was teleoperated, how many attempts were filmed, whether the scene was staged and whether it was sped up.
Taken together, this is a field making real progress on a problem that is harder than it looks from a highlight reel, and the way robots learn now depends far more on data and evaluation than on any single breakthrough. The same questions are worth asking of any autonomy claim, which is why the staged definitions in our explainer on levels of self-driving are a useful habit of mind, and why a consumer cleaning robot remains a narrow, carefully constrained product rather than a general helper.
Common questions
Can a robot really learn a task by watching once? Not in the way the phrase suggests. Policies are trained on many recorded demonstrations, often pooled across machines, and a model that adapts quickly to a new task is drawing on that prior training rather than learning from a single viewing.
What is imitation learning? Training a model to map sensor input to actions using recorded human demonstrations, usually collected by teleoperating the robot through the task many times.
Why train in simulation at all if it is inaccurate? Because it is cheap and fast, and because the inaccuracy can be turned into an advantage. Domain randomization varies rendering and physics so widely that the real world becomes one more variation the model has effectively seen.
What is the hardest part of robot manipulation? Contact. Force, friction and deformable objects are poorly simulated and poorly represented in demonstration data, so folding cloth is harder than moving a rigid box.
How should I judge a robot demonstration video? Ask whether it was teleoperated, how many takes were recorded, whether the environment was arranged for the shot, whether playback was sped up, and whether the same system has been evaluated on a published benchmark.
Sources and further reading
Where the figures and rules above come from, so you can check them:
- 22 robots, 21 institutions, 527 categories and 160,266 tasks, and positive transfer between platforms: Open X-Embodiment and RT-X, arXiv
- Diffusion Policy, 12 tasks across 4 benchmarks and the 46.9 percent average improvement: arXiv
- pi-0: flow matching on a pre-trained vision-language model, platforms and evaluated tasks: arXiv
- Domain randomization and the 1.5 cm localisation accuracy transferred from simulation: Tobin et al., arXiv
- Generality, interactivity and dexterity, and the generalisation benchmark claim: Google DeepMind
- Performance metrics and test methods for robotic grasping and assembly: NIST Intelligent Systems Division
Photo credit: Controlling robotic arm ESA15746126 by European Space Agency, CC BY-SA 3.0 igo, via Wikimedia Commons.
Join the discussion