Imitation learning trains a robot to replicate demonstrated behavior. The tighter the correspondence between the demonstration data and the robot's own sensory input, the better the learned policy generalizes. Egocentric data achieves this correspondence directly — the training images look like what the robot will see.
In behavior cloning, the policy maps observation to action at each timestep. With ego video, the observation is a first-person image plus depth, and the action is the recorded hand or arm motion. There is no viewpoint transformation, no camera calibration mismatch, and no occlusion of the end effector. This reduces compounding error, the primary failure mode of behavior cloning at deployment time.
For diffusion policy and transformer-based action models, large-scale egocentric datasets provide the diversity needed for generalization. Thousands of hours of first-person demonstrations across environments, objects, and operators give these models the distributional coverage to handle novel situations at test time.
Behavior cloningDiffusion policyTransformer action modelsVisuomotor policies