Change the representation.
Predict the state directly.
Is tracking the hidden stickers itself expensive, or does the difficulty come from learning state transitions through video? We replace the video representation with the colors of the 12 visible stickers. Stickers still disappear from view, and the model must predict how their colors return.
Video space
Wan VAE
Vector space
encoding
A small state Transformer tracks these transitions with high accuracy over 20 moves, using much less compute. The hidden-state task remains, but removing video generation makes it substantially easier in this controlled setting.
| State model | Training PF-days | Final-state accuracy | Full-trajectory accuracy |
|---|---|---|---|
| AR · 2.8M | ≈ 0.008 | 94.01% | 92.81% |
| Bidirectional · 2.8M | ≈ 0.0004 | 99.86% | 99.85% |
Final-state accuracy requires all visible sticker colors to be correct after move 20. Full-trajectory accuracy requires every predicted state to be correct. These are separate state-only experiments.


