Visuo-Acoustic Hand Pose and Contact Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915456461307904 |
|---|---|
| author | Mao, Yuemin Yoo, Uksang Yao, Yunchao Syed, Shahram Najam Bondi, Luca Francis, Jonathan Oh, Jean Ichnowski, Jeffrey |
| author_facet | Mao, Yuemin Yoo, Uksang Yao, Yunchao Syed, Shahram Najam Bondi, Luca Francis, Jonathan Oh, Jean Ichnowski, Jeffrey |
| contents | Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains challenging due to visual occlusion, subtle contact cues, limitations in vision-only sensing, and the lack of accessible and flexible tactile sensing. We therefore introduce VibeMesh, a novel wearable system that fuses vision with active acoustic sensing for dense, per-vertex hand contact and pose estimation. VibeMesh integrates a bone-conduction speaker and sparse piezoelectric microphones, distributed on a human hand, emitting structured acoustic signals and capturing their propagation to infer changes induced by contact. To interpret these cross-modal signals, we propose a graph-based attention network that processes synchronized audio spectra and RGB-D-derived hand meshes to predict contact with high spatial resolution. We contribute: (i) a lightweight, non-intrusive visuo-acoustic sensing platform; (ii) a cross-modal graph network for joint pose and contact inference; (iii) a dataset of synchronized RGB-D, acoustic, and ground-truth contact annotations across diverse manipulation scenarios; and (iv) empirical results showing that VibeMesh outperforms vision-only baselines in accuracy and robustness, particularly in occluded or static-contact settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_00852 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Visuo-Acoustic Hand Pose and Contact Estimation Mao, Yuemin Yoo, Uksang Yao, Yunchao Syed, Shahram Najam Bondi, Luca Francis, Jonathan Oh, Jean Ichnowski, Jeffrey Human-Computer Interaction Computer Vision and Pattern Recognition Machine Learning Robotics Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains challenging due to visual occlusion, subtle contact cues, limitations in vision-only sensing, and the lack of accessible and flexible tactile sensing. We therefore introduce VibeMesh, a novel wearable system that fuses vision with active acoustic sensing for dense, per-vertex hand contact and pose estimation. VibeMesh integrates a bone-conduction speaker and sparse piezoelectric microphones, distributed on a human hand, emitting structured acoustic signals and capturing their propagation to infer changes induced by contact. To interpret these cross-modal signals, we propose a graph-based attention network that processes synchronized audio spectra and RGB-D-derived hand meshes to predict contact with high spatial resolution. We contribute: (i) a lightweight, non-intrusive visuo-acoustic sensing platform; (ii) a cross-modal graph network for joint pose and contact inference; (iii) a dataset of synchronized RGB-D, acoustic, and ground-truth contact annotations across diverse manipulation scenarios; and (iv) empirical results showing that VibeMesh outperforms vision-only baselines in accuracy and robustness, particularly in occluded or static-contact settings. |
| title | Visuo-Acoustic Hand Pose and Contact Estimation |
| topic | Human-Computer Interaction Computer Vision and Pattern Recognition Machine Learning Robotics |
| url | https://arxiv.org/abs/2508.00852 |