Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jinzhou, Wu, Tianhao, Zhang, Jiyao, Chen, Zeyuan, Jin, Haotian, Wu, Mingdong, Shen, Yujun, Yang, Yaodong, Dong, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916853556707328
author Li, Jinzhou
Wu, Tianhao
Zhang, Jiyao
Chen, Zeyuan
Jin, Haotian
Wu, Mingdong
Shen, Yujun
Yang, Yaodong
Dong, Hao
author_facet Li, Jinzhou
Wu, Tianhao
Zhang, Jiyao
Chen, Zeyuan
Jin, Haotian
Wu, Mingdong
Shen, Yujun
Yang, Yaodong
Dong, Hao
contents Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain comprehensively fused features but often ignore the fact that each modality requires different levels of attention at different manipulation stages. To address this, we propose a force-guided attention fusion module that adaptively adjusts the weights of visual and tactile features without human labeling. We also introduce a self-supervised future force prediction auxiliary task to reinforce the tactile modality, improve data imbalance, and encourage proper adjustment. Our method achieves an average success rate of 93% across three fine-grained, contactrich tasks in real-world experiments. Further analysis shows that our policy appropriately adjusts attention to each modality at different manipulation stages. The videos can be viewed at https://adaptac-dex.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13982
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation
Li, Jinzhou
Wu, Tianhao
Zhang, Jiyao
Chen, Zeyuan
Jin, Haotian
Wu, Mingdong
Shen, Yujun
Yang, Yaodong
Dong, Hao
Robotics
Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain comprehensively fused features but often ignore the fact that each modality requires different levels of attention at different manipulation stages. To address this, we propose a force-guided attention fusion module that adaptively adjusts the weights of visual and tactile features without human labeling. We also introduce a self-supervised future force prediction auxiliary task to reinforce the tactile modality, improve data imbalance, and encourage proper adjustment. Our method achieves an average success rate of 93% across three fine-grained, contactrich tasks in real-world experiments. Further analysis shows that our policy appropriately adjusts attention to each modality at different manipulation stages. The videos can be viewed at https://adaptac-dex.github.io/.
title Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation
topic Robotics
url https://arxiv.org/abs/2505.13982