Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mane, Atharv Mahesh, Weerakoon, Dulanga, Subbaraju, Vigneshwaran, Sen, Sougata, Sarma, Sanjay E., Misra, Archan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908317221126144
author Mane, Atharv Mahesh
Weerakoon, Dulanga
Subbaraju, Vigneshwaran
Sen, Sougata
Sarma, Sanjay E.
Misra, Archan
author_facet Mane, Atharv Mahesh
Weerakoon, Dulanga
Subbaraju, Vigneshwaran
Sen, Sougata
Sarma, Sanjay E.
Misra, Archan
contents 3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. To address this gap, we introduce a data augmentation framework-Imputer, and use it to curate a new benchmark dataset-ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. We also propose Ges3ViG, a novel model for 3D-ERU that achieves ~30% improvement in accuracy as compared to other 3D-ERU models and ~9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/Ges3ViG.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
Mane, Atharv Mahesh
Weerakoon, Dulanga
Subbaraju, Vigneshwaran
Sen, Sougata
Sarma, Sanjay E.
Misra, Archan
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. To address this gap, we introduce a data augmentation framework-Imputer, and use it to curate a new benchmark dataset-ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. We also propose Ges3ViG, a novel model for 3D-ERU that achieves ~30% improvement in accuracy as compared to other 3D-ERU models and ~9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/Ges3ViG.
title Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2504.09623