Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuan, Tianyuan, Dong, Zibin, Liu, Yicheng, Zhao, Hang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
von: Liu, Chenyv, et al.
Veröffentlicht: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
von: Wang, Linbo, et al.
Veröffentlicht: (2026)
von: Wang, Linbo, et al.
Veröffentlicht: (2026)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
von: Liu, Yulong, et al.
Veröffentlicht: (2024)
von: Liu, Yulong, et al.
Veröffentlicht: (2024)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
von: Peng, Jierui, et al.
Veröffentlicht: (2025)
von: Peng, Jierui, et al.
Veröffentlicht: (2025)
Sparse Imagination for Efficient Visual World Model Planning
von: Chun, Junha, et al.
Veröffentlicht: (2025)
von: Chun, Junha, et al.
Veröffentlicht: (2025)
Capsule Networks Do Not Need to Model Everything
von: Renzulli, Riccardo, et al.
Veröffentlicht: (2022)
von: Renzulli, Riccardo, et al.
Veröffentlicht: (2022)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
von: Han, Jianhua, et al.
Veröffentlicht: (2025)
von: Han, Jianhua, et al.
Veröffentlicht: (2025)
WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
Learning Vision-Language-Action World Models for Autonomous Driving
von: Wang, Guoqing, et al.
Veröffentlicht: (2026)
von: Wang, Guoqing, et al.
Veröffentlicht: (2026)
Do We Need Large VLMs for Spotting Soccer Actions?
von: Chakraborty, Ritabrata, et al.
Veröffentlicht: (2025)
von: Chakraborty, Ritabrata, et al.
Veröffentlicht: (2025)
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
von: Xu, Hang, et al.
Veröffentlicht: (2025)
von: Xu, Hang, et al.
Veröffentlicht: (2025)
WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving
von: Xu, Mingwang, et al.
Veröffentlicht: (2025)
von: Xu, Mingwang, et al.
Veröffentlicht: (2025)
Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
MIND: Benchmarking Memory Consistency and Action Control in World Models
von: Ye, Yixuan, et al.
Veröffentlicht: (2026)
von: Ye, Yixuan, et al.
Veröffentlicht: (2026)
Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling
von: Jha, Saurav, et al.
Veröffentlicht: (2025)
von: Jha, Saurav, et al.
Veröffentlicht: (2025)
Learning Latent Action World Models In The Wild
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
von: Garrido, Quentin, et al.
Veröffentlicht: (2026)
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
von: Fang, Jianjie, et al.
Veröffentlicht: (2026)
von: Fang, Jianjie, et al.
Veröffentlicht: (2026)
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
von: Shi, Chen, et al.
Veröffentlicht: (2026)
von: Shi, Chen, et al.
Veröffentlicht: (2026)
Interpretable and Sparse Linear Attention with Decoupled Membership-Subspace Modeling via MCR2 Objective
von: Liu, Tianyuan, et al.
Veröffentlicht: (2026)
von: Liu, Tianyuan, et al.
Veröffentlicht: (2026)
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2024)
Ideal Registration? Segmentation is All You Need
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
von: Chen, Xiang, et al.
Veröffentlicht: (2025)
SFMViT: SlowFast Meet ViT in Chaotic World
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
von: Guo, Jun, et al.
Veröffentlicht: (2026)
von: Guo, Jun, et al.
Veröffentlicht: (2026)
FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
von: Xu, Hang, et al.
Veröffentlicht: (2025)
von: Xu, Hang, et al.
Veröffentlicht: (2025)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
von: Perincherry, Akhil, et al.
Veröffentlicht: (2025)
DiLA: Disentangled Latent Action World Models
von: Zhang, Tianqiu, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqiu, et al.
Veröffentlicht: (2026)
Do We Really Need a Large Number of Visual Prompts?
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
von: Sun, Zhengqi, et al.
Veröffentlicht: (2026)
von: Sun, Zhengqi, et al.
Veröffentlicht: (2026)
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
von: Huang, YiJie, et al.
Veröffentlicht: (2026)
von: Huang, YiJie, et al.
Veröffentlicht: (2026)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
von: Liu, Yicheng, et al.
Veröffentlicht: (2025)
DualFast: Dual-Speedup Framework for Fast Sampling of Diffusion Models
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
RenderWorld: World Model with Self-Supervised 3D Label
von: Yan, Ziyang, et al.
Veröffentlicht: (2024)
von: Yan, Ziyang, et al.
Veröffentlicht: (2024)
Pandora: Towards General World Model with Natural Language Actions and Video States
von: Xiang, Jiannan, et al.
Veröffentlicht: (2024)
von: Xiang, Jiannan, et al.
Veröffentlicht: (2024)
OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
von: Wu, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Wu, Zhiqiang, et al.
Veröffentlicht: (2025)
Anatomy Might Be All You Need: Forecasting What to Do During Surgery
von: Sarwin, Gary, et al.
Veröffentlicht: (2025)
von: Sarwin, Gary, et al.
Veröffentlicht: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
von: Chen, Yuxin, et al.
Veröffentlicht: (2025)
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
von: Samuel, Dvir, et al.
Veröffentlicht: (2026)
von: Samuel, Dvir, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
von: Liu, Chenyv, et al.
Veröffentlicht: (2026) -
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
von: Yuan, Tianyuan, et al.
Veröffentlicht: (2025) -
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
von: Wang, Linbo, et al.
Veröffentlicht: (2026) -
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
von: Yu, Shoubin, et al.
Veröffentlicht: (2026) -
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
von: Liu, Yulong, et al.
Veröffentlicht: (2024)