EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Binzhu, Qiu, Shi, Zhang, Sicheng, Wang, Yinqiao, Xu, Hao, Naseer, Muzammal, Fu, Chi-Wing, Heng, Pheng-Ann
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918309069324288
author Xie, Binzhu
Qiu, Shi
Zhang, Sicheng
Wang, Yinqiao
Xu, Hao
Naseer, Muzammal
Fu, Chi-Wing
Heng, Pheng-Ann
author_facet Xie, Binzhu
Qiu, Shi
Zhang, Sicheng
Wang, Yinqiao
Xu, Hao
Naseer, Muzammal
Fu, Chi-Wing
Heng, Pheng-Ann
contents Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they often struggle in unseen contexts. We present EgoHandICL, the first in-context learning (ICL) framework for 3D hand reconstruction that improves semantic alignment, visual consistency, and robustness under challenging egocentric conditions. EgoHandICL introduces complementary exemplar retrieval guided by vision-language models (VLMs), an ICL-tailored tokenizer for multimodal context, and a masked autoencoder (MAE)-based architecture trained with hand-guided geometric and perceptual objectives. Experiments on ARCTIC and EgoExo4D show consistent gains over state-of-the-art methods. We also demonstrate real-world generalization and improve EgoVLM hand-object interaction reasoning by using reconstructed hands as visual prompts. Code and data: https://github.com/Nicous20/EgoHandICL
format Preprint
id arxiv_https___arxiv_org_abs_2601_19850
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
Xie, Binzhu
Qiu, Shi
Zhang, Sicheng
Wang, Yinqiao
Xu, Hao
Naseer, Muzammal
Fu, Chi-Wing
Heng, Pheng-Ann
Computer Vision and Pattern Recognition
Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they often struggle in unseen contexts. We present EgoHandICL, the first in-context learning (ICL) framework for 3D hand reconstruction that improves semantic alignment, visual consistency, and robustness under challenging egocentric conditions. EgoHandICL introduces complementary exemplar retrieval guided by vision-language models (VLMs), an ICL-tailored tokenizer for multimodal context, and a masked autoencoder (MAE)-based architecture trained with hand-guided geometric and perceptual objectives. Experiments on ARCTIC and EgoExo4D show consistent gains over state-of-the-art methods. We also demonstrate real-world generalization and improve EgoVLM hand-object interaction reasoning by using reconstructed hands as visual prompts. Code and data: https://github.com/Nicous20/EgoHandICL
title EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.19850