WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Weilun, Fan, Guoxin, Qin, Haotong, Wu, Mingqiang, Li, Yuqi, Li, Xiangqi, An, Zhulin, Huang, Libo, Wang, Dingrui, Liao, Longlong, Magno, Michele, Xu, Yongjun, Yang, Chuanguang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916071477346304
author Feng, Weilun
Fan, Guoxin
Qin, Haotong
Wu, Mingqiang
Li, Yuqi
Li, Xiangqi
An, Zhulin
Huang, Libo
Wang, Dingrui
Liao, Longlong
Magno, Michele
Xu, Yongjun
Yang, Chuanguang
author_facet Feng, Weilun
Fan, Guoxin
Qin, Haotong
Wu, Mingqiang
Li, Yuqi
Li, Xiangqi
An, Zhulin
Huang, Libo
Wang, Dingrui
Liao, Longlong
Magno, Michele
Xu, Yongjun
Yang, Chuanguang
contents Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffusion transfer poorly to world models due to two world-model-specific obstacles: \emph{token heterogeneity} from multi-modal coupling and spatial variation, and \emph{non-uniform temporal dynamics} where a small set of hard tokens drives error growth, making uniform skipping either unstable or overly conservative. We propose \textbf{WorldCache}, a caching framework tailored to diffusion world models. We introduce \textit{Curvature-guided Heterogeneous Token Prediction}, which uses a physics-grounded curvature score to estimate token predictability and applies a Hermite-guided damped predictor for chaotic tokens with abrupt direction changes. We also design \textit{Chaotic-prioritized Adaptive Skipping}, which accumulates a curvature-normalized, dimensionless drift signal and recomputes only when bottleneck tokens begin to drift. Experiments on diffusion world models show that WorldCache delivers up to \textbf{3.7$\times$} end-to-end speedups while maintaining \textbf{98\%} rollout quality, demonstrating the vast advantages and practicality of WorldCache in resource-constrained scenarios. Our code is released in https://github.com/FofGofx/WorldCache.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06331
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
Feng, Weilun
Fan, Guoxin
Qin, Haotong
Wu, Mingqiang
Li, Yuqi
Li, Xiangqi
An, Zhulin
Huang, Libo
Wang, Dingrui
Liao, Longlong
Magno, Michele
Xu, Yongjun
Yang, Chuanguang
Computer Vision and Pattern Recognition
Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactive use and long-horizon rollouts. While feature caching can accelerate inference without training, we find that policies designed for single-modal diffusion transfer poorly to world models due to two world-model-specific obstacles: \emph{token heterogeneity} from multi-modal coupling and spatial variation, and \emph{non-uniform temporal dynamics} where a small set of hard tokens drives error growth, making uniform skipping either unstable or overly conservative. We propose \textbf{WorldCache}, a caching framework tailored to diffusion world models. We introduce \textit{Curvature-guided Heterogeneous Token Prediction}, which uses a physics-grounded curvature score to estimate token predictability and applies a Hermite-guided damped predictor for chaotic tokens with abrupt direction changes. We also design \textit{Chaotic-prioritized Adaptive Skipping}, which accumulates a curvature-normalized, dimensionless drift signal and recomputes only when bottleneck tokens begin to drift. Experiments on diffusion world models show that WorldCache delivers up to \textbf{3.7$\times$} end-to-end speedups while maintaining \textbf{98\%} rollout quality, demonstrating the vast advantages and practicality of WorldCache in resource-constrained scenarios. Our code is released in https://github.com/FofGofx/WorldCache.
title WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.06331