EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866918309069324288 |
|---|---|
| author | Xie, Binzhu Qiu, Shi Zhang, Sicheng Wang, Yinqiao Xu, Hao Naseer, Muzammal Fu, Chi-Wing Heng, Pheng-Ann |
| author_facet | Xie, Binzhu Qiu, Shi Zhang, Sicheng Wang, Yinqiao Xu, Hao Naseer, Muzammal Fu, Chi-Wing Heng, Pheng-Ann |
| contents | Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they often struggle in unseen contexts. We present EgoHandICL, the first in-context learning (ICL) framework for 3D hand reconstruction that improves semantic alignment, visual consistency, and robustness under challenging egocentric conditions. EgoHandICL introduces complementary exemplar retrieval guided by vision-language models (VLMs), an ICL-tailored tokenizer for multimodal context, and a masked autoencoder (MAE)-based architecture trained with hand-guided geometric and perceptual objectives. Experiments on ARCTIC and EgoExo4D show consistent gains over state-of-the-art methods. We also demonstrate real-world generalization and improve EgoVLM hand-object interaction reasoning by using reconstructed hands as visual prompts. Code and data: https://github.com/Nicous20/EgoHandICL |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_19850 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning Xie, Binzhu Qiu, Shi Zhang, Sicheng Wang, Yinqiao Xu, Hao Naseer, Muzammal Fu, Chi-Wing Heng, Pheng-Ann Computer Vision and Pattern Recognition Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues by scaling training data or adding auxiliary cues, but they often struggle in unseen contexts. We present EgoHandICL, the first in-context learning (ICL) framework for 3D hand reconstruction that improves semantic alignment, visual consistency, and robustness under challenging egocentric conditions. EgoHandICL introduces complementary exemplar retrieval guided by vision-language models (VLMs), an ICL-tailored tokenizer for multimodal context, and a masked autoencoder (MAE)-based architecture trained with hand-guided geometric and perceptual objectives. Experiments on ARCTIC and EgoExo4D show consistent gains over state-of-the-art methods. We also demonstrate real-world generalization and improve EgoVLM hand-object interaction reasoning by using reconstructed hands as visual prompts. Code and data: https://github.com/Nicous20/EgoHandICL |
| title | EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.19850 |