FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yicheng, Zhang, Shiduo, Dong, Zibin, Ye, Baijun, Yuan, Tianyuan, Yu, Xiaopeng, Yin, Linqi, Lu, Chenhao, Shi, Junhao, Yu, Luca Jiang-Tao, Zheng, Liangtao, Jiang, Tao, Gong, Jingjing, Qiu, Xipeng, Zhao, Hang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915660771098624
author Liu, Yicheng
Zhang, Shiduo
Dong, Zibin
Ye, Baijun
Yuan, Tianyuan
Yu, Xiaopeng
Yin, Linqi
Lu, Chenhao
Shi, Junhao
Yu, Luca Jiang-Tao
Zheng, Liangtao
Jiang, Tao
Gong, Jingjing
Qiu, Xipeng
Zhao, Hang
author_facet Liu, Yicheng
Zhang, Shiduo
Dong, Zibin
Ye, Baijun
Yuan, Tianyuan
Yu, Xiaopeng
Yin, Linqi
Lu, Chenhao
Shi, Junhao
Yu, Luca Jiang-Tao
Zheng, Liangtao
Jiang, Tao
Gong, Jingjing
Qiu, Xipeng
Zhao, Hang
contents Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce FASTer, a unified framework for efficient and generalizable robot learning that integrates a learnable tokenizer with an autoregressive policy built upon it. FASTerVQ encodes action chunks as single-channel images, capturing global spatio-temporal dependencies while maintaining a high compression ratio. FASTerVLA builds on this tokenizer with block-wise autoregressive decoding and a lightweight action expert, achieving both faster inference and higher task performance. Extensive experiments across simulated and real-world benchmarks show that FASTerVQ delivers superior reconstruction quality, high token utilization, and strong cross-task and cross-embodiment generalization, while FASTerVLA further improves overall capability, surpassing previous state-of-the-art VLA models in both inference speed and task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
Liu, Yicheng
Zhang, Shiduo
Dong, Zibin
Ye, Baijun
Yuan, Tianyuan
Yu, Xiaopeng
Yin, Linqi
Lu, Chenhao
Shi, Junhao
Yu, Luca Jiang-Tao
Zheng, Liangtao
Jiang, Tao
Gong, Jingjing
Qiu, Xipeng
Zhao, Hang
Computer Vision and Pattern Recognition
Robotics
Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce FASTer, a unified framework for efficient and generalizable robot learning that integrates a learnable tokenizer with an autoregressive policy built upon it. FASTerVQ encodes action chunks as single-channel images, capturing global spatio-temporal dependencies while maintaining a high compression ratio. FASTerVLA builds on this tokenizer with block-wise autoregressive decoding and a lightweight action expert, achieving both faster inference and higher task performance. Extensive experiments across simulated and real-world benchmarks show that FASTerVQ delivers superior reconstruction quality, high token utilization, and strong cross-task and cross-embodiment generalization, while FASTerVLA further improves overall capability, surpassing previous state-of-the-art VLA models in both inference speed and task performance.
title FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.04952