Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Fuente:
arXiv
Guardado en:
| Autores principales: | Yuan, Tianyuan, Dong, Zibin, Liu, Yicheng, Zhao, Hang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
por: Liu, Chenyv, et al.
Publicado: (2026)
por: Liu, Chenyv, et al.
Publicado: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
por: Yuan, Tianyuan, et al.
Publicado: (2025)
por: Yuan, Tianyuan, et al.
Publicado: (2025)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
por: Wang, Linbo, et al.
Publicado: (2026)
por: Wang, Linbo, et al.
Publicado: (2026)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
por: Yu, Shoubin, et al.
Publicado: (2026)
por: Yu, Shoubin, et al.
Publicado: (2026)
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
por: Liu, Yulong, et al.
Publicado: (2024)
por: Liu, Yulong, et al.
Publicado: (2024)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
por: Peng, Jierui, et al.
Publicado: (2025)
por: Peng, Jierui, et al.
Publicado: (2025)
Sparse Imagination for Efficient Visual World Model Planning
por: Chun, Junha, et al.
Publicado: (2025)
por: Chun, Junha, et al.
Publicado: (2025)
Capsule Networks Do Not Need to Model Everything
por: Renzulli, Riccardo, et al.
Publicado: (2022)
por: Renzulli, Riccardo, et al.
Publicado: (2022)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
por: Han, Jianhua, et al.
Publicado: (2025)
por: Han, Jianhua, et al.
Publicado: (2025)
WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
por: Xu, Yifang, et al.
Publicado: (2025)
por: Xu, Yifang, et al.
Publicado: (2025)
Learning Vision-Language-Action World Models for Autonomous Driving
por: Wang, Guoqing, et al.
Publicado: (2026)
por: Wang, Guoqing, et al.
Publicado: (2026)
Do We Need Large VLMs for Spotting Soccer Actions?
por: Chakraborty, Ritabrata, et al.
Publicado: (2025)
por: Chakraborty, Ritabrata, et al.
Publicado: (2025)
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
por: Xu, Hang, et al.
Publicado: (2025)
por: Xu, Hang, et al.
Publicado: (2025)
WAM-Diff: A Masked Diffusion VLA Framework with MoE and Online Reinforcement Learning for Autonomous Driving
por: Xu, Mingwang, et al.
Publicado: (2025)
por: Xu, Mingwang, et al.
Publicado: (2025)
Schrödinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
por: He, Yu, et al.
Publicado: (2025)
por: He, Yu, et al.
Publicado: (2025)
MIND: Benchmarking Memory Consistency and Action Control in World Models
por: Ye, Yixuan, et al.
Publicado: (2026)
por: Ye, Yixuan, et al.
Publicado: (2026)
Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling
por: Jha, Saurav, et al.
Publicado: (2025)
por: Jha, Saurav, et al.
Publicado: (2025)
Learning Latent Action World Models In The Wild
por: Garrido, Quentin, et al.
Publicado: (2026)
por: Garrido, Quentin, et al.
Publicado: (2026)
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
por: Fang, Jianjie, et al.
Publicado: (2026)
por: Fang, Jianjie, et al.
Publicado: (2026)
DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving
por: Shi, Chen, et al.
Publicado: (2026)
por: Shi, Chen, et al.
Publicado: (2026)
Interpretable and Sparse Linear Attention with Decoupled Membership-Subspace Modeling via MCR2 Objective
por: Liu, Tianyuan, et al.
Publicado: (2026)
por: Liu, Tianyuan, et al.
Publicado: (2026)
Do We Really Need a Complex Agent System? Distill Embodied Agent into a Single Model
por: Zhao, Zhonghan, et al.
Publicado: (2024)
por: Zhao, Zhonghan, et al.
Publicado: (2024)
Ideal Registration? Segmentation is All You Need
por: Chen, Xiang, et al.
Publicado: (2025)
por: Chen, Xiang, et al.
Publicado: (2025)
SFMViT: SlowFast Meet ViT in Chaotic World
por: Lin, Jiaying, et al.
Publicado: (2024)
por: Lin, Jiaying, et al.
Publicado: (2024)
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
por: Guo, Jun, et al.
Publicado: (2026)
por: Guo, Jun, et al.
Publicado: (2026)
FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
por: Xu, Hang, et al.
Publicado: (2025)
por: Xu, Hang, et al.
Publicado: (2025)
Do Visual Imaginations Improve Vision-and-Language Navigation Agents?
por: Perincherry, Akhil, et al.
Publicado: (2025)
por: Perincherry, Akhil, et al.
Publicado: (2025)
DiLA: Disentangled Latent Action World Models
por: Zhang, Tianqiu, et al.
Publicado: (2026)
por: Zhang, Tianqiu, et al.
Publicado: (2026)
Do We Really Need a Large Number of Visual Prompts?
por: Kim, Youngeun, et al.
Publicado: (2023)
por: Kim, Youngeun, et al.
Publicado: (2023)
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
por: Sun, Zhengqi, et al.
Publicado: (2026)
por: Sun, Zhengqi, et al.
Publicado: (2026)
How Many Visual Tokens Do Multimodal Language Models Need? Scaling Visual Token Pruning with F^3A
por: Huang, YiJie, et al.
Publicado: (2026)
por: Huang, YiJie, et al.
Publicado: (2026)
FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
por: Liu, Yicheng, et al.
Publicado: (2025)
por: Liu, Yicheng, et al.
Publicado: (2025)
DualFast: Dual-Speedup Framework for Fast Sampling of Diffusion Models
por: Yu, Hu, et al.
Publicado: (2025)
por: Yu, Hu, et al.
Publicado: (2025)
RenderWorld: World Model with Self-Supervised 3D Label
por: Yan, Ziyang, et al.
Publicado: (2024)
por: Yan, Ziyang, et al.
Publicado: (2024)
Pandora: Towards General World Model with Natural Language Actions and Video States
por: Xiang, Jiannan, et al.
Publicado: (2024)
por: Xiang, Jiannan, et al.
Publicado: (2024)
OMGSR: You Only Need One Mid-timestep Guidance for Real-World Image Super-Resolution
por: Wu, Zhiqiang, et al.
Publicado: (2025)
por: Wu, Zhiqiang, et al.
Publicado: (2025)
Anatomy Might Be All You Need: Forecasting What to Do During Surgery
por: Sarwin, Gary, et al.
Publicado: (2025)
por: Sarwin, Gary, et al.
Publicado: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
por: Huang, Haiduo, et al.
Publicado: (2025)
por: Huang, Haiduo, et al.
Publicado: (2025)
Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
por: Chen, Yuxin, et al.
Publicado: (2025)
por: Chen, Yuxin, et al.
Publicado: (2025)
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
por: Samuel, Dvir, et al.
Publicado: (2026)
por: Samuel, Dvir, et al.
Publicado: (2026)
Ejemplares similares
-
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
por: Liu, Chenyv, et al.
Publicado: (2026) -
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
por: Yuan, Tianyuan, et al.
Publicado: (2025) -
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
por: Wang, Linbo, et al.
Publicado: (2026) -
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
por: Yu, Shoubin, et al.
Publicado: (2026) -
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
por: Liu, Yulong, et al.
Publicado: (2024)