A predicted future can show what success might look like without specifying the exact state a controller must reach. Robots need compact, measurable goals—such as an object’s final position and orientation—and continuous feedback about the real scene. A plausible future image is evidence for planning, not yet an executable instruction.
A picture of success leaves important questions unanswered
Suppose a world model generates a future frame in which a red cube sits on a blue cube. The image communicates the intended outcome to a person. A robot controller still needs to know which object is the target, its position in three-dimensional space, its orientation, the allowable error, and how those quantities change as the gripper moves.
Pixels contain some of that information indirectly. Extracting it after generation can introduce ambiguity: the cube may look correctly placed while floating slightly above the surface, appearing at the wrong depth, or occupying a pose the gripper cannot produce.
Prediction and execution use different representations
The preprint Predicted Futures Are Not Enough, released 7 October 2026 by Tzu-Yu Chuang and colleagues, makes the interface explicit. Its world model predicts how a manipulation scene evolves, but it also produces an entity-level goal: an object-centred target pose expressed in SE(3), the standard mathematical description of three-dimensional position and rotation.
A separate pose-native executor consumes that fixed goal and uses online object-pose feedback for closed-loop control. The world model does not need to rerun at every control step. Across five reported manipulation tasks, the pipeline reaches a mean success rate of 79.69%. On a Franka robot, the paper reports 73.33% zero-shot success for its nominal StackCube setting, 66.67% with distractors, and 75.00% for PickPlate with a target unseen during policy training.
Those results belong to the authors’ tasks and setup; they do not show that any predicted video can be converted reliably into robot action.
Closed-loop control keeps asking the world what actually happened
An executable goal does not prescribe every motor command in advance. The controller observes the object again, compares its current pose with the target, acts, and repeats. This loop matters because grasps slip, cameras are noisy, and contact changes the scene.
That is the difference between a destination and a route. The goal says where the object should end. Feedback determines the next correction from where it is now.
Somatic practice offers a familiar analogy. Imagining the shape of a balance does not specify every muscular adjustment needed to enter it. The practitioner continuously senses weight, support, and drift. The imagined form orients action; it does not replace perception.
Why not let the video model control everything directly?
An end-to-end system may learn a mapping from predicted frames to action, but then failures are difficult to locate. Was the future visually wrong? Was the object pose extracted incorrectly? Was the goal unreachable? Did control fail after a good goal?
An explicit prediction-to-execution interface creates a diagnostic boundary. The paper tests controlled translation errors to measure how execution degrades as the goal becomes inaccurate. That makes goal quality inspectable instead of treating it as incidental post-processing.
The same principle appeared in Somatic-AI Lab’s analysis of the planning limits of latent world models: imagination is most useful when its horizon and handoff to observation are visible. The October frontier report likewise argues that representations should be judged by the action-relevant distinctions they preserve.
The practical takeaway
When a system claims to plan through generated futures, ask what crosses the boundary into control. Is it a picture, a latent vector, a trajectory, an object pose, or a verified subgoal? Then ask how the system detects that the real world has diverged.
A convincing imagined future may help a robot choose an outcome. Reliable action begins when that future is translated into a goal the controller can measure, pursue, and revise.
References
- Chuang, T.-Y., Chang, C.-H., Lee, Y.-H., Chen, Y.-T., Sun, M., & Yang, Y. (2026). Predicted futures are not enough: Learning executable goals for robot manipulation. arXiv. https://arxiv.org/abs/2610.09309