Scaling Video Generation
for Reasoning: At What Cost?

A convincing video can still predict the wrong state.
How far does scaling take us?

Weihang Guo1Xiaoyu Wu2Yifei Wang1Niloofar Mireshghallah2Lydia E. Kavraki1

1 Rice University2 Carnegie Mellon University

The right shape. The wrong state.

Scrub through every frame of a generated sequence.

1B model · 3M training videos

Loading the saved rollout…

Ground truth VAE reconstruction
Ground-truth cube after action 6.
Generated video AR-k1 · EMA weights
Generated cube after action 6.

Drag through all 81 frames, including the initial image, or select a move to inspect its final state. A red outline marks an incorrect action-boundary frame. Random switches among 20 saved rollouts across model sizes, training checkpoints, and action sequences. Ground truth is reconstructed through the same frozen VAE as the predictions.

01 / A controlled test

Nine moves. Three visible faces.
One correct future.

Start with a solved 2 × 2 × 2 Rubik’s Cube and prescribe nine turns. A fixed camera sees three faces. As the cube turns, stickers disappear from view and later return in different positions. Correct video prediction requires keeping track of these hidden state changes.

The solved initial state and action sequence determine every future configuration. This gives us exact targets for evaluating state prediction, while leaving the model free to learn its own representation.

Input
Initial image + all nine prescribed moves
Prediction
80 frames (3.3 s) of future video
Frame accuracy
Fraction of post-action frames with all 12 visible stickers correct.
Full-trajectory accuracy
Fraction of episodes with all 12 visible stickers correct after every one of the nine moves.

The benchmark measures prediction under prescribed actions. It does not ask the model to find a solution to a scrambled cube.

02 / Measure the state

Lower loss does not reliably
identify better reasoning.

Validation flow MSE improves with scale, but similar losses can accompany very different state accuracies across models. A model’s progress in fitting video latents therefore needs to be checked against the states it actually generates.

Each point pairs validation loss and free-running accuracy from the same EMA checkpoint. Within a model, lower MSE generally tracks learning progress. Across models, it is not a common scale of reasoning capability.

03 / The cost of reliability

Scaling helps.
Reliable trajectories remain costly.

Smaller autoregressive models perform better when compute is limited. Larger models reach higher accuracy with more training. Yet improvements at individual steps do not guarantee that the entire sequence stays correct: every move is another opportunity for a state error.

All results use the same 100 held-out episodes and generated history. Frame accuracy requires all 12 visible stickers to be correct at a step; full-trajectory accuracy requires this after every one of the nine moves.

What would near-perfect reliability cost?

We fit the observed compute frontier and extend the curves toward 98% accuracy.

Compute frontier and power-law extrapolations toward 98 percent accuracy, comparing two fitting windows for AR frame accuracy, AR full-trajectory accuracy, and bidirectional frame accuracy.

Open full-size plot

Gray points show evaluated checkpoints; colored points mark the middle and final thirds of the observed frontier. Solid lines show fits within each window; dashed lines extend those fits. Shading marks the observed compute range, and diamonds mark the projected 98% intersections. Bidirectional full-trajectory accuracy is omitted because it is zero at every evaluated checkpoint.

04 / Locate the difficulty

What makes hidden-state
prediction hard?

A

Change the representation.
Predict the state directly.

Is tracking the hidden stickers itself expensive, or does the difficulty come from learning state transitions through video? We replace the video representation with the colors of the 12 visible stickers. Stickers still disappear from view, and the model must predict how their colors return.

Same kind of conditioning: initial observation + prescribed action prompt

Video space

Video framesRendered observations
Frozen
Wan VAE
Continuous latents16 × 32 × 32 per latent frame
Prediction modelVideo DiTAR or bidirectional

Vector space

Visible sticker colorsSimulator labels
One-hot
encoding
Categorical vectors12 stickers × 6 colors
Prediction modelState TransformerAR or bidirectional · 2.8M
A comparison of training representations. The vector model predicts visible sticker colors directly, while the video model predicts continuous video latents.

A small state Transformer tracks these transitions with high accuracy over 20 moves, using much less compute. The hidden-state task remains, but removing video generation makes it substantially easier in this controlled setting.

Direct state prediction · 20 actions · 10,000 held-out episodes
State modelTraining PF-daysFinal-state accuracyFull-trajectory accuracy
AR · 2.8M≈ 0.00894.01%92.81%
Bidirectional · 2.8M≈ 0.000499.86%99.85%

Final-state accuracy requires all visible sticker colors to be correct after move 20. Full-trajectory accuracy requires every predicted state to be correct. These are separate state-only experiments.

B

Change what is visible.
Keep every face in view.

Next, keep video prediction and reveal the hidden faces. A complementary camera shows the other three faces, so the model can observe all 24 stickers. The physical transitions stay the same, while the additional view changes both state visibility and the video the model must predict.

The three panels below follow the same actions. Compare the three-face prediction with its ground truth, then inspect the six-face prediction at exactly the same frame.

3-face ground truthSimulator
3-face prediction12 visible stickers

6-face prediction24 visible stickers

Loading the paired observation experiment…

A paired rollout from the separate 270M AR-k4 observation experiment, with 7.68M training video presentations per model. Drag through every frame; select a move to compare scored states. Red outlines mark incorrect predictions at action boundaries. The three-face criterion checks 12 stickers; the six-face criterion checks all 24.

Across the evaluation set, full observation substantially improves video state accuracy, especially for autoregressive generation. Together, these controls point to the difficulty of maintaining state information through video generation when part of the world is hidden.

05 / Symbolic guidance

Predict the state.
Use it to guide the video.

A shared Transformer predicts sticker-state distributions from the action prompt and available video history. Those predictions then condition a separate video pass. Training combines state supervision with the video flow objective; generation uses the model’s own state predictions.

Original paper workflow: video history and the action prompt condition separate state and video passes with shared Transformer weights. Predicted sticker-state probabilities feed back to the video pass. Training combines state cross-entropy and video flow loss.

Open full-size figure

Symbolic-guidance training, from the paper. Predicted sticker-state probabilities condition the video pass; both passes share Transformer weights. Training combines state cross-entropy with the video flow objective.

20M AR-k1 models, each trained on 3M videos and evaluated on 100 paired episodes using EMA weights. Both guided variants supervise all 24 stickers. Front12 feeds back the 12 visible-sticker distributions; Full24 also feeds back the hidden stickers.

Explicit state supervision and predicted-state feedback improve video state accuracy at the same training-data budget. The result suggests that learning representations of state change can complement scaling.

The open question

What state should
a reasoning model learn?

In the cube, we know what must be remembered: sticker colors, their positions, and how actions change them. Symbolic guidance helps when that state is defined for the model. The broader challenge is to discover useful state abstractions from experience—and make them work across objects, tasks, and environments.

Our symbolic targets are discrete, while the model feeds back probability distributions over those targets. The result leaves open how a general reasoning system should represent state: as discrete entities and relations, continuous latent variables, or a combination of both.

Approaches such as JEPA explore prediction in learned representation spaces. Whether those representations capture the state needed for reliable reasoning, or whether a different structure is required, remains a research question.

Can models learn for themselves which aspects of the world must be preserved, how actions transform them, and how to use that knowledge in unfamiliar situations?

Citation

@misc{guo2026scalingvideogenerationreasoning,
  title={Scaling Video Generation for Reasoning: At What Cost?},
  author={Weihang Guo and Xiaoyu Wu and Yifei Wang and
          Niloofar Mireshghallah and Lydia E. Kavraki},
  year={2026},
  eprint={2609.36599},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.36599}
}