I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Qi, Yu, Gu, Lipeng, Chen, Honghua, Nan, Liangliang, Wei, Mingqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908410760396800
author Qi, Yu
Gu, Lipeng
Chen, Honghua
Nan, Liangliang
Wei, Mingqiang
author_facet Qi, Yu
Gu, Lipeng
Chen, Honghua
Nan, Liangliang
Wei, Mingqiang
contents Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from transcription errors due to accents, background noise, and varying speech rates, limiting the applicability of existing 3DVG methods. To address these challenges, we propose \textbf{SpeechRefer}, a novel 3DVG framework designed to enhance performance in the presence of noisy and ambiguous speech-to-text transcriptions. SpeechRefer integrates seamlessly with xisting 3DVG models and introduces two key innovations. First, the Speech Complementary Module captures acoustic similarities between phonetically related words and highlights subtle distinctions, generating complementary proposal scores from the speech signal. This reduces dependence on potentially erroneous transcriptions. Second, the Contrastive Complementary Module employs contrastive learning to align erroneous text features with corresponding speech features, ensuring robust performance even when transcription errors dominate. Extensive experiments on the SpeechRefer and peechNr3D datasets demonstrate that SpeechRefer improves the performance of existing 3DVG methods by a large margin, which highlights SpeechRefer's potential to bridge the gap between noisy speech inputs and reliable 3DVG, enabling more intuitive and practical multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs
Qi, Yu
Gu, Lipeng
Chen, Honghua
Nan, Liangliang
Wei, Mingqiang
Computer Vision and Pattern Recognition
Existing 3D visual grounding methods rely on precise text prompts to locate objects within 3D scenes. Speech, as a natural and intuitive modality, offers a promising alternative. Real-world speech inputs, however, often suffer from transcription errors due to accents, background noise, and varying speech rates, limiting the applicability of existing 3DVG methods. To address these challenges, we propose \textbf{SpeechRefer}, a novel 3DVG framework designed to enhance performance in the presence of noisy and ambiguous speech-to-text transcriptions. SpeechRefer integrates seamlessly with xisting 3DVG models and introduces two key innovations. First, the Speech Complementary Module captures acoustic similarities between phonetically related words and highlights subtle distinctions, generating complementary proposal scores from the speech signal. This reduces dependence on potentially erroneous transcriptions. Second, the Contrastive Complementary Module employs contrastive learning to align erroneous text features with corresponding speech features, ensuring robust performance even when transcription errors dominate. Extensive experiments on the SpeechRefer and peechNr3D datasets demonstrate that SpeechRefer improves the performance of existing 3DVG methods by a large margin, which highlights SpeechRefer's potential to bridge the gap between noisy speech inputs and reliable 3DVG, enabling more intuitive and practical multimodal systems.
title I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.14495