Video world models — AI systems that generate navigable, spatially coherent video from a single starting image — have a fundamental memory problem that makes them unreliable for the robotics training ...