TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Zongzheng, Xu, Haobo, Yang, Zhuo, Yue, Chenghao, Lin, Zehao, Gao, Huan-ang, Wang, Ziwei, Zhao, Hao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908528120168448
author Zhang, Zongzheng
Xu, Haobo
Yang, Zhuo
Yue, Chenghao
Lin, Zehao
Gao, Huan-ang
Wang, Ziwei
Zhao, Hao
author_facet Zhang, Zongzheng
Xu, Haobo
Yang, Zhuo
Yue, Chenghao
Lin, Zehao
Gao, Huan-ang
Wang, Ziwei
Zhao, Hao
contents Many robotic manipulation tasks require sensing and responding to force signals such as torque to assess whether the task has been successfully completed and to enable closed-loop control. However, current Vision-Language-Action (VLA) models lack the ability to integrate such subtle physical feedback. In this work, we explore Torque-aware VLA models, aiming to bridge this gap by systematically studying the design space for incorporating torque signals into existing VLA architectures. We identify and evaluate several strategies, leading to three key findings. First, introducing torque adapters into the decoder consistently outperforms inserting them into the encoder.Third, inspired by joint prediction and planning paradigms in autonomous driving, we propose predicting torque as an auxiliary output, which further improves performance. This strategy encourages the model to build a physically grounded internal representation of interaction dynamics. Extensive quantitative and qualitative experiments across contact-rich manipulation benchmarks validate our findings.
format Preprint
id arxiv_https___arxiv_org_abs_2509_07962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
Zhang, Zongzheng
Xu, Haobo
Yang, Zhuo
Yue, Chenghao
Lin, Zehao
Gao, Huan-ang
Wang, Ziwei
Zhao, Hao
Robotics
Many robotic manipulation tasks require sensing and responding to force signals such as torque to assess whether the task has been successfully completed and to enable closed-loop control. However, current Vision-Language-Action (VLA) models lack the ability to integrate such subtle physical feedback. In this work, we explore Torque-aware VLA models, aiming to bridge this gap by systematically studying the design space for incorporating torque signals into existing VLA architectures. We identify and evaluate several strategies, leading to three key findings. First, introducing torque adapters into the decoder consistently outperforms inserting them into the encoder.Third, inspired by joint prediction and planning paradigms in autonomous driving, we propose predicting torque as an auxiliary output, which further improves performance. This strategy encourages the model to build a physically grounded internal representation of interaction dynamics. Extensive quantitative and qualitative experiments across contact-rich manipulation benchmarks validate our findings.
title TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2509.07962