SGGNet$^2$: Speech-Scene Graph Grounding Network for Speech-guided Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Dohyun, Kim, Yeseung, Jang, Jaehwi, Song, Minjae, Choi, Woojin, Park, Daehyung
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911838000644096
author Kim, Dohyun
Kim, Yeseung
Jang, Jaehwi
Song, Minjae
Choi, Woojin
Park, Daehyung
author_facet Kim, Dohyun
Kim, Yeseung
Jang, Jaehwi
Song, Minjae
Choi, Woojin
Park, Daehyung
contents The spoken language serves as an accessible and efficient interface, enabling non-experts and disabled users to interact with complex assistant robots. However, accurately grounding language utterances gives a significant challenge due to the acoustic variability in speakers' voices and environmental noise. In this work, we propose a novel speech-scene graph grounding network (SGGNet$^2$) that robustly grounds spoken utterances by leveraging the acoustic similarity between correctly recognized and misrecognized words obtained from automatic speech recognition (ASR) systems. To incorporate the acoustic similarity, we extend our previous grounding model, the scene-graph-based grounding network (SGGNet), with the ASR model from NVIDIA NeMo. We accomplish this by feeding the latent vector of speech pronunciations into the BERT-based grounding network within SGGNet. We evaluate the effectiveness of using latent vectors of speech commands in grounding through qualitative and quantitative studies. We also demonstrate the capability of SGGNet$^2$ in a speech-based navigation task using a real quadruped robot, RBQ-3, from Rainbow Robotics.
format Preprint
id arxiv_https___arxiv_org_abs_2307_07468
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SGGNet$^2$: Speech-Scene Graph Grounding Network for Speech-guided Navigation
Kim, Dohyun
Kim, Yeseung
Jang, Jaehwi
Song, Minjae
Choi, Woojin
Park, Daehyung
Robotics
The spoken language serves as an accessible and efficient interface, enabling non-experts and disabled users to interact with complex assistant robots. However, accurately grounding language utterances gives a significant challenge due to the acoustic variability in speakers' voices and environmental noise. In this work, we propose a novel speech-scene graph grounding network (SGGNet$^2$) that robustly grounds spoken utterances by leveraging the acoustic similarity between correctly recognized and misrecognized words obtained from automatic speech recognition (ASR) systems. To incorporate the acoustic similarity, we extend our previous grounding model, the scene-graph-based grounding network (SGGNet), with the ASR model from NVIDIA NeMo. We accomplish this by feeding the latent vector of speech pronunciations into the BERT-based grounding network within SGGNet. We evaluate the effectiveness of using latent vectors of speech commands in grounding through qualitative and quantitative studies. We also demonstrate the capability of SGGNet$^2$ in a speech-based navigation task using a real quadruped robot, RBQ-3, from Rainbow Robotics.
title SGGNet$^2$: Speech-Scene Graph Grounding Network for Speech-guided Navigation
topic Robotics
url https://arxiv.org/abs/2307.07468