What 30,000 Hours of Ego-centric Video Does Not Teach
Preprint, 2026
How far does scaling ego-centric human video take world models? We train on 30,000 hours of video spanning 1,000+ scene types and find that scaling brings agent modeling near saturation but leaves object dynamics far behind. Closing this gap depends on how models are trained, not on data alone.










