VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Jianxin, Ma, Kevin Yuchen, Hao, Ce, Shou, Mike Zheng, Soh, Harold
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912506819117056
author Bi, Jianxin
Ma, Kevin Yuchen
Hao, Ce
Shou, Mike Zheng
Soh, Harold
author_facet Bi, Jianxin
Ma, Kevin Yuchen
Hao, Ce
Shou, Mike Zheng
Soh, Harold
contents Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the ability to interpret and use tactile signals, limiting their effectiveness in contact-rich tasks. Incorporating tactile feedback into these systems is challenging due to the absence of large multi-modal datasets. We present VLA-Touch, an approach that enhances generalist robot policies with tactile sensing \emph{without fine-tuning} the base VLA. Our method introduces two key innovations: (1) a pipeline that leverages a pretrained tactile-language model that provides semantic tactile feedback for high-level task planning, and (2) a diffusion-based controller that refines VLA-generated actions with tactile signals for contact-rich manipulation. Through real-world experiments, we demonstrate that our dual-level integration of tactile feedback improves task planning efficiency while enhancing execution precision. Code is open-sourced at \href{https://github.com/jxbi1010/VLA-Touch}{this URL}.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17294
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
Bi, Jianxin
Ma, Kevin Yuchen
Hao, Ce
Shou, Mike Zheng
Soh, Harold
Robotics
Machine Learning
Tactile feedback is generally recognized to be crucial for effective interaction with the physical world. However, state-of-the-art Vision-Language-Action (VLA) models lack the ability to interpret and use tactile signals, limiting their effectiveness in contact-rich tasks. Incorporating tactile feedback into these systems is challenging due to the absence of large multi-modal datasets. We present VLA-Touch, an approach that enhances generalist robot policies with tactile sensing \emph{without fine-tuning} the base VLA. Our method introduces two key innovations: (1) a pipeline that leverages a pretrained tactile-language model that provides semantic tactile feedback for high-level task planning, and (2) a diffusion-based controller that refines VLA-generated actions with tactile signals for contact-rich manipulation. Through real-world experiments, we demonstrate that our dual-level integration of tactile feedback improves task planning efficiency while enhancing execution precision. Code is open-sourced at \href{https://github.com/jxbi1010/VLA-Touch}{this URL}.
title VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
topic Robotics
Machine Learning
url https://arxiv.org/abs/2507.17294