TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Yuzhe, Lin, Pei, Li, Wanlin, Li, Daohan, Li, Jiajun, Jiang, Jiaming, Xiao, Chenxi, Jiao, Ziyuan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911410637766656
author Huang, Yuzhe
Lin, Pei
Li, Wanlin
Li, Daohan
Li, Jiajun
Jiang, Jiaming
Xiao, Chenxi
Jiao, Ziyuan
author_facet Huang, Yuzhe
Lin, Pei
Li, Wanlin
Li, Daohan
Li, Jiajun
Jiang, Jiaming
Xiao, Chenxi
Jiao, Ziyuan
contents Vision-Language-Action (VLA) models have recently emerged as powerful generalists for robotic manipulation. However, due to their predominant reliance on visual modalities, they fundamentally lack the physical intuition required for contact-rich tasks that require precise force regulation and physical reasoning. Existing attempts to incorporate vision-based tactile sensing into VLA models typically treat tactile inputs as auxiliary visual textures, thereby overlooking the underlying correlation between surface deformation and interaction dynamics. To bridge this gap, we propose a paradigm shift from tactile-vision alignment to tactile-force alignment. Here, we introduce TaF-VLA, a framework that explicitly grounds high-dimensional tactile observations in physical interaction forces. To facilitate this, we develop an automated tactile-force data acquisition device and curate the TaF-Dataset, comprising over 10 million synchronized tactile observations, 6-axis force/torque, and matrix force map. To align sequential tactile observations with interaction forces, the central component of our approach is the Tactile-Force Adapter (TaF-Adapter), a tactile sensor encoder that extracts discretized latent information for encoding tactile observations. This mechanism ensures that the learned representations capture history-dependent, noise-insensitive physical dynamics rather than static visual textures. Finally, we integrate this force-aligned encoder into a VLA backbone. Extensive real-world experiments demonstrate that TaF-VLA policy significantly outperforms state-of-the-art tactile-vision-aligned and vision-only baselines on contact-rich tasks, verifying its ability to achieve robust, force-aware manipulation through cross-modal physical reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20321
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
Huang, Yuzhe
Lin, Pei
Li, Wanlin
Li, Daohan
Li, Jiajun
Jiang, Jiaming
Xiao, Chenxi
Jiao, Ziyuan
Robotics
Vision-Language-Action (VLA) models have recently emerged as powerful generalists for robotic manipulation. However, due to their predominant reliance on visual modalities, they fundamentally lack the physical intuition required for contact-rich tasks that require precise force regulation and physical reasoning. Existing attempts to incorporate vision-based tactile sensing into VLA models typically treat tactile inputs as auxiliary visual textures, thereby overlooking the underlying correlation between surface deformation and interaction dynamics. To bridge this gap, we propose a paradigm shift from tactile-vision alignment to tactile-force alignment. Here, we introduce TaF-VLA, a framework that explicitly grounds high-dimensional tactile observations in physical interaction forces. To facilitate this, we develop an automated tactile-force data acquisition device and curate the TaF-Dataset, comprising over 10 million synchronized tactile observations, 6-axis force/torque, and matrix force map. To align sequential tactile observations with interaction forces, the central component of our approach is the Tactile-Force Adapter (TaF-Adapter), a tactile sensor encoder that extracts discretized latent information for encoding tactile observations. This mechanism ensures that the learned representations capture history-dependent, noise-insensitive physical dynamics rather than static visual textures. Finally, we integrate this force-aligned encoder into a VLA backbone. Extensive real-world experiments demonstrate that TaF-VLA policy significantly outperforms state-of-the-art tactile-vision-aligned and vision-only baselines on contact-rich tasks, verifying its ability to achieve robust, force-aware manipulation through cross-modal physical reasoning.
title TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
topic Robotics
url https://arxiv.org/abs/2601.20321