FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915660771098624 |
|---|---|
| author | Liu, Yicheng Zhang, Shiduo Dong, Zibin Ye, Baijun Yuan, Tianyuan Yu, Xiaopeng Yin, Linqi Lu, Chenhao Shi, Junhao Yu, Luca Jiang-Tao Zheng, Liangtao Jiang, Tao Gong, Jingjing Qiu, Xipeng Zhao, Hang |
| author_facet | Liu, Yicheng Zhang, Shiduo Dong, Zibin Ye, Baijun Yuan, Tianyuan Yu, Xiaopeng Yin, Linqi Lu, Chenhao Shi, Junhao Yu, Luca Jiang-Tao Zheng, Liangtao Jiang, Tao Gong, Jingjing Qiu, Xipeng Zhao, Hang |
| contents | Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce FASTer, a unified framework for efficient and generalizable robot learning that integrates a learnable tokenizer with an autoregressive policy built upon it. FASTerVQ encodes action chunks as single-channel images, capturing global spatio-temporal dependencies while maintaining a high compression ratio. FASTerVLA builds on this tokenizer with block-wise autoregressive decoding and a lightweight action expert, achieving both faster inference and higher task performance. Extensive experiments across simulated and real-world benchmarks show that FASTerVQ delivers superior reconstruction quality, high token utilization, and strong cross-task and cross-embodiment generalization, while FASTerVLA further improves overall capability, surpassing previous state-of-the-art VLA models in both inference speed and task performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_04952 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization Liu, Yicheng Zhang, Shiduo Dong, Zibin Ye, Baijun Yuan, Tianyuan Yu, Xiaopeng Yin, Linqi Lu, Chenhao Shi, Junhao Yu, Luca Jiang-Tao Zheng, Liangtao Jiang, Tao Gong, Jingjing Qiu, Xipeng Zhao, Hang Computer Vision and Pattern Recognition Robotics Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce FASTer, a unified framework for efficient and generalizable robot learning that integrates a learnable tokenizer with an autoregressive policy built upon it. FASTerVQ encodes action chunks as single-channel images, capturing global spatio-temporal dependencies while maintaining a high compression ratio. FASTerVLA builds on this tokenizer with block-wise autoregressive decoding and a lightweight action expert, achieving both faster inference and higher task performance. Extensive experiments across simulated and real-world benchmarks show that FASTerVQ delivers superior reconstruction quality, high token utilization, and strong cross-task and cross-embodiment generalization, while FASTerVLA further improves overall capability, surpassing previous state-of-the-art VLA models in both inference speed and task performance. |
| title | FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2512.04952 |