Visuo-Acoustic Hand Pose and Contact Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Yuemin, Yoo, Uksang, Yao, Yunchao, Syed, Shahram Najam, Bondi, Luca, Francis, Jonathan, Oh, Jean, Ichnowski, Jeffrey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915456461307904
author Mao, Yuemin
Yoo, Uksang
Yao, Yunchao
Syed, Shahram Najam
Bondi, Luca
Francis, Jonathan
Oh, Jean
Ichnowski, Jeffrey
author_facet Mao, Yuemin
Yoo, Uksang
Yao, Yunchao
Syed, Shahram Najam
Bondi, Luca
Francis, Jonathan
Oh, Jean
Ichnowski, Jeffrey
contents Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains challenging due to visual occlusion, subtle contact cues, limitations in vision-only sensing, and the lack of accessible and flexible tactile sensing. We therefore introduce VibeMesh, a novel wearable system that fuses vision with active acoustic sensing for dense, per-vertex hand contact and pose estimation. VibeMesh integrates a bone-conduction speaker and sparse piezoelectric microphones, distributed on a human hand, emitting structured acoustic signals and capturing their propagation to infer changes induced by contact. To interpret these cross-modal signals, we propose a graph-based attention network that processes synchronized audio spectra and RGB-D-derived hand meshes to predict contact with high spatial resolution. We contribute: (i) a lightweight, non-intrusive visuo-acoustic sensing platform; (ii) a cross-modal graph network for joint pose and contact inference; (iii) a dataset of synchronized RGB-D, acoustic, and ground-truth contact annotations across diverse manipulation scenarios; and (iv) empirical results showing that VibeMesh outperforms vision-only baselines in accuracy and robustness, particularly in occluded or static-contact settings.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visuo-Acoustic Hand Pose and Contact Estimation
Mao, Yuemin
Yoo, Uksang
Yao, Yunchao
Syed, Shahram Najam
Bondi, Luca
Francis, Jonathan
Oh, Jean
Ichnowski, Jeffrey
Human-Computer Interaction
Computer Vision and Pattern Recognition
Machine Learning
Robotics
Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains challenging due to visual occlusion, subtle contact cues, limitations in vision-only sensing, and the lack of accessible and flexible tactile sensing. We therefore introduce VibeMesh, a novel wearable system that fuses vision with active acoustic sensing for dense, per-vertex hand contact and pose estimation. VibeMesh integrates a bone-conduction speaker and sparse piezoelectric microphones, distributed on a human hand, emitting structured acoustic signals and capturing their propagation to infer changes induced by contact. To interpret these cross-modal signals, we propose a graph-based attention network that processes synchronized audio spectra and RGB-D-derived hand meshes to predict contact with high spatial resolution. We contribute: (i) a lightweight, non-intrusive visuo-acoustic sensing platform; (ii) a cross-modal graph network for joint pose and contact inference; (iii) a dataset of synchronized RGB-D, acoustic, and ground-truth contact annotations across diverse manipulation scenarios; and (iv) empirical results showing that VibeMesh outperforms vision-only baselines in accuracy and robustness, particularly in occluded or static-contact settings.
title Visuo-Acoustic Hand Pose and Contact Estimation
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2508.00852