Why first person video may matter for robot learning[D]
4/10I can see why first-person video might help a robot model, but not because the robot can copy a human hand. The joints, reach, timing, and control space are all different. What may transfer is the sequence of visual attention: which object enters view, what changes before contact, and where the actor looks next. LingBot-VLA 2.0 (arXiv:2607.06403) uses first-person data alongside robot trajectories. A useful ablation separates visual prediction from robot control. First-person pretraining might improve next-state prediction without improving task success, which would still tell us where the in
