FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yicheng, Zhang, Shiduo, Dong, Zibin, Ye, Baijun, Yuan, Tianyuan, Yu, Xiaopeng, Yin, Linqi, Lu, Chenhao, Shi, Junhao, Yu, Luca Jiang-Tao, Zheng, Liangtao, Jiang, Tao, Gong, Jingjing, Qiu, Xipeng, Zhao, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ActionCodec: What Makes for Good Action Tokenizers
by: Dong, Zibin, et al.
Published: (2026)
by: Dong, Zibin, et al.
Published: (2026)
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
by: Yuan, Tianyuan, et al.
Published: (2025)
by: Yuan, Tianyuan, et al.
Published: (2025)
Fast-WAM: Do World Action Models Need Test-time Future Imagination?
by: Yuan, Tianyuan, et al.
Published: (2026)
by: Yuan, Tianyuan, et al.
Published: (2026)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)
by: Fei, Senyu, et al.
Published: (2025)
World Action Models: The Next Frontier in Embodied AI
by: Wang, Siyin, et al.
Published: (2026)
by: Wang, Siyin, et al.
Published: (2026)
World-aware Planning Narratives Enhance Large Vision-Language Model Planner
by: Shi, Junhao, et al.
Published: (2025)
by: Shi, Junhao, et al.
Published: (2025)
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
by: Zhang, Shiduo, et al.
Published: (2024)
by: Zhang, Shiduo, et al.
Published: (2024)
CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization
by: Chen, Yitong, et al.
Published: (2026)
by: Chen, Yitong, et al.
Published: (2026)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
by: Peng, Jierui, et al.
Published: (2025)
by: Peng, Jierui, et al.
Published: (2025)
FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning
by: Guo, Hang, et al.
Published: (2025)
by: Guo, Hang, et al.
Published: (2025)
Beyond Attention Magnitude: Leveraging Inter-layer Rank Consistency for Efficient Vision-Language-Action Models
by: Liu, Peiju, et al.
Published: (2026)
by: Liu, Peiju, et al.
Published: (2026)
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
by: Dai, Yuntao, et al.
Published: (2025)
by: Dai, Yuntao, et al.
Published: (2025)
FASTer: Focal Token Acquiring-and-Scaling Transformer for Long-term 3D Object Detection
by: Dang, Chenxu, et al.
Published: (2025)
by: Dang, Chenxu, et al.
Published: (2025)
Drive My Way: Preference Alignment of Vision-Language-Action Model for Personalized Driving
by: Wang, Zehao, et al.
Published: (2026)
by: Wang, Zehao, et al.
Published: (2026)
FAST: Efficient Action Tokenization for Vision-Language-Action Models
by: Pertsch, Karl, et al.
Published: (2025)
by: Pertsch, Karl, et al.
Published: (2025)
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models
by: Hu, Yutong, et al.
Published: (2026)
by: Hu, Yutong, et al.
Published: (2026)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
by: Qiu, Yicheng, et al.
Published: (2026)
by: Qiu, Yicheng, et al.
Published: (2026)
Galaxea Open-World Dataset and G0 Dual-System VLA Model
by: Jiang, Tao, et al.
Published: (2025)
by: Jiang, Tao, et al.
Published: (2025)
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
by: Fei, Zhaoye, et al.
Published: (2025)
by: Fei, Zhaoye, et al.
Published: (2025)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
by: Xu, Zonghuan, et al.
Published: (2025)
by: Xu, Zonghuan, et al.
Published: (2025)
Enhancing Large Language Models for Mobility Analytics with Semantic Location Tokenization
by: Chen, Yile, et al.
Published: (2025)
by: Chen, Yile, et al.
Published: (2025)
WorldVLA: Towards Autoregressive Action World Model
by: Cen, Jun, et al.
Published: (2025)
by: Cen, Jun, et al.
Published: (2025)
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
by: Xu, Siyu, et al.
Published: (2025)
by: Xu, Siyu, et al.
Published: (2025)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
by: Yu, Wenda, et al.
Published: (2026)
by: Yu, Wenda, et al.
Published: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
by: Zhong, Yifan, et al.
Published: (2025)
by: Zhong, Yifan, et al.
Published: (2025)
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation
by: Zhang, Wenbo, et al.
Published: (2025)
by: Zhang, Wenbo, et al.
Published: (2025)
Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
by: Liu, Ting, et al.
Published: (2024)
by: Liu, Ting, et al.
Published: (2024)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
by: Zhao, Baining, et al.
Published: (2026)
by: Zhao, Baining, et al.
Published: (2026)
Actor-Critic for Continuous Action Chunks: A Reinforcement Learning Framework for Long-Horizon Robotic Manipulation with Sparse Reward
by: Yang, Jiarui, et al.
Published: (2025)
by: Yang, Jiarui, et al.
Published: (2025)
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation
by: Liu, Guoshan, et al.
Published: (2026)
by: Liu, Guoshan, et al.
Published: (2026)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
by: Zheng, Weicheng, et al.
Published: (2026)
by: Zheng, Weicheng, et al.
Published: (2026)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
by: Wang, Siyin, et al.
Published: (2025)
by: Wang, Siyin, et al.
Published: (2025)
A Survey on Efficient Vision-Language-Action Models
by: Yu, Zhaoshu, et al.
Published: (2025)
by: Yu, Zhaoshu, et al.
Published: (2025)
AMP: Autoregressive Motion Prediction Revisited with Next Token Prediction for Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2024)
by: Jia, Xiaosong, et al.
Published: (2024)
Similar Items
-
ActionCodec: What Makes for Good Action Tokenizers
by: Dong, Zibin, et al.
Published: (2026) -
DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
by: Yuan, Tianyuan, et al.
Published: (2025) -
Fast-WAM: Do World Action Models Need Test-time Future Imagination?
by: Yuan, Tianyuan, et al.
Published: (2026) -
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025) -
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
by: Fei, Senyu, et al.
Published: (2025)