Teaser

Why generate pixels when the policy operates on VLM tokens?

Token-World teaser figure
Teaser figure placeholderassets/teaser.png
Token-World simulates directly in compact VLM visual-token space. Unlike RGB-based world models that generate future images and re-encode them into policy inputs, Token-World models action-conditioned dynamics directly in compact VLM tokens. This design avoids intermediate RGB generation, reduces simulator step time to 0.359 s, and yields more reliable closed-loop policy evaluation (r = 0.794 vs. 0.583 for Ctrl-World).

Method Overview

Token-World method overview
Method figure placeholderassets/method_overview.png
Token-World overview. (a) Compact Token Construction. Qwen3-VL visual tokens are compressed into a compact latent space. (b) Dynamics Modeling. A spatiotemporal Transformer models action-conditioned dynamics in compact token space. (c) Training Objective. The model is trained with diffusion forcing and weighted flow matching. (d) Autoregressive Rollout. Future compact tokens are generated with a sliding temporal window and recurrent state.

Qualitative Results

Quantitative Results

Open-loop comparison of VLM feature fidelity and policy-action consistency on RoboTwin and real-world Franka tasks, measured by cosine similarity and NMSE
Open-loop prediction figure placeholderassets/open_loop_prediction.png
Open-loop prediction quality. Token-World achieves the strongest VLM feature fidelity and policy-action consistency across both RoboTwin and real-world Franka tasks, with higher cosine similarity and lower NMSE than RGB-based world models.
Policy evaluation scatter plots comparing Token-World and Ctrl-World against reference success rates, alongside average simulation step times
Policy evaluation and efficiency figure placeholderassets/policy_evaluation_efficiency.png
Policy evaluation and efficiency. Token-World’s predicted outcomes align more closely with actual policy success rates (r = 0.794 vs. 0.583 for Ctrl-World), while also achieving the fastest simulation speed at 0.359 s per step.

Ablations

Ablation studies comparing VLM and RGB VAE representations, raw and compressed VLM features, and compact latent dimensions of 8, 16, 32, and 48
Ablation figure placeholderassets/ablation.png
Ablation studies. (a) VLM-based compact states improve both feature fidelity and action consistency over RGB VAE representations. (b) Compressing raw VLM features substantially improves dynamics prediction. (c) A 16-D compact representation provides the best overall trade-off between feature and action consistency.