AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Ju, Moritz, Niko, Huang, Yiteng, Xie, Ruiming, Sun, Ming, Fuegen, Christian, Seide, Frank
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917570519498752
author Lin, Ju
Moritz, Niko
Huang, Yiteng
Xie, Ruiming
Sun, Ming
Fuegen, Christian
Seide, Frank
author_facet Lin, Ju
Moritz, Niko
Huang, Yiteng
Xie, Ruiming
Sun, Ming
Fuegen, Christian
Seide, Frank
contents Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise. When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses. This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28\% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10411
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
Lin, Ju
Moritz, Niko
Huang, Yiteng
Xie, Ruiming
Sun, Ming
Fuegen, Christian
Seide, Frank
Audio and Speech Processing
Sound
Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise. When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses. This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28\% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.
title AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2401.10411