A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Kerssies, Tommie, Berton, Gabriele, He, Ju, Yu, Qihang, Ma, Wufei, de Geus, Daan, Dubbelman, Gijs, Chen, Liang-Chieh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
by: Englert, Brunó B., et al.
Published: (2024)
by: Englert, Brunó B., et al.
Published: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
What is the Added Value of UDA in the VFM Era?
by: Englert, Brunó B., et al.
Published: (2025)
by: Englert, Brunó B., et al.
Published: (2025)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
by: de Geus, Daan, et al.
Published: (2024)
by: de Geus, Daan, et al.
Published: (2024)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
by: Norouzi, Narges, et al.
Published: (2026)
by: Norouzi, Narges, et al.
Published: (2026)
Your ViT is Secretly an Image Segmentation Model
by: Kerssies, Tommie, et al.
Published: (2025)
by: Kerssies, Tommie, et al.
Published: (2025)
Simplifying Traffic Anomaly Detection with Video Foundation Models
by: Orlova, Svetlana, et al.
Published: (2025)
by: Orlova, Svetlana, et al.
Published: (2025)
An Image is Worth 32 Tokens for Reconstruction and Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
by: He, Ju, et al.
Published: (2025)
by: He, Ju, et al.
Published: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
MaskBit: Embedding-free Image Generation via Bit Tokens
by: Weber, Mark, et al.
Published: (2024)
by: Weber, Mark, et al.
Published: (2024)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
by: Orlova, Svetlana, et al.
Published: (2026)
by: Orlova, Svetlana, et al.
Published: (2026)
Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models
by: Gu, Jing, et al.
Published: (2026)
by: Gu, Jing, et al.
Published: (2026)
One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy
by: Tang, Zuojin, et al.
Published: (2026)
by: Tang, Zuojin, et al.
Published: (2026)
Randomized Autoregressive Visual Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation
by: Englert, Brunó B., et al.
Published: (2025)
by: Englert, Brunó B., et al.
Published: (2025)
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
by: Vu, Tuan-Hung, et al.
Published: (2024)
by: Vu, Tuan-Hung, et al.
Published: (2024)
MegaLoc: One Retrieval to Place Them All
by: Berton, Gabriele, et al.
Published: (2025)
by: Berton, Gabriele, et al.
Published: (2025)
REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View
by: Chandrasekaran, Kavin, et al.
Published: (2026)
by: Chandrasekaran, Kavin, et al.
Published: (2026)
A Resource Efficient Fusion Network for Object Detection in Bird's-Eye View using Camera and Raw Radar Data
by: Chandrasekaran, Kavin, et al.
Published: (2024)
by: Chandrasekaran, Kavin, et al.
Published: (2024)
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
by: Zhou, Qihang, et al.
Published: (2025)
by: Zhou, Qihang, et al.
Published: (2025)
Frequency-Aware Flow Matching for High-Quality Image Generation
by: Ren, Sucheng, et al.
Published: (2026)
by: Ren, Sucheng, et al.
Published: (2026)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
by: Liu, Qihao, et al.
Published: (2024)
by: Liu, Qihao, et al.
Published: (2024)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
by: Wang, Feng, et al.
Published: (2025)
by: Wang, Feng, et al.
Published: (2025)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
A Creative Agent is Worth a 64-Token Template
by: Shi, Ruixiao, et al.
Published: (2026)
by: Shi, Ruixiao, et al.
Published: (2026)
DONUT: A Decoder-Only Model for Trajectory Prediction
by: Knoche, Markus, et al.
Published: (2025)
by: Knoche, Markus, et al.
Published: (2025)
Autoregressive Image Generation with Masked Bit Modeling
by: Yu, Qihang, et al.
Published: (2026)
by: Yu, Qihang, et al.
Published: (2026)
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
Token Bottleneck: One Token to Remember Dynamics
by: Kim, Taekyung, et al.
Published: (2025)
by: Kim, Taekyung, et al.
Published: (2025)
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
by: Zhang, Zheyu, et al.
Published: (2026)
by: Zhang, Zheyu, et al.
Published: (2026)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
by: Nezhurina, Marianna, et al.
Published: (2025)
by: Nezhurina, Marianna, et al.
Published: (2025)
Similar Items
-
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024) -
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
by: Kerssies, Tommie, et al.
Published: (2024) -
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
by: Englert, Brunó B., et al.
Published: (2024) -
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024) -
What is the Added Value of UDA in the VFM Era?
by: Englert, Brunó B., et al.
Published: (2025)