Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Yufeng, Huang, Yiteng, Xu, Yong, Wan, Li, Shon, Suwon, Liu, Yang, Fan, Yifeng, Yang, Zhaojun, Siohan, Olivier, Liu, Yue, Sun, Ming, Metze, Florian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911160793563136
author Yang, Yufeng
Huang, Yiteng
Xu, Yong
Wan, Li
Shon, Suwon
Liu, Yang
Fan, Yifeng
Yang, Zhaojun
Siohan, Olivier
Liu, Yue
Sun, Ming
Metze, Florian
author_facet Yang, Yufeng
Huang, Yiteng
Xu, Yong
Wan, Li
Shon, Suwon
Liu, Yang
Fan, Yifeng
Yang, Zhaojun
Siohan, Olivier
Liu, Yue
Sun, Ming
Metze, Florian
contents With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-computer interfaces. However, in real environments, interference from side-talk speech remains a significant challenge to WSR and may cause accumulated errors for downstream tasks such as natural language processing. In this work, we introduce a novel multi-channel differential automatic speech recognition (ASR) method for robust WSR on smart glasses. The proposed system takes differential inputs from different frontends that complement each other to improve the robustness of WSR, including a beamformer, microphone selection, and a lightweight side-talk detection model. Evaluations on both simulated and real datasets demonstrate that the proposed system outperforms the traditional approach, achieving up to an 18.0% relative reduction in word error rate.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
Yang, Yufeng
Huang, Yiteng
Xu, Yong
Wan, Li
Shon, Suwon
Liu, Yang
Fan, Yifeng
Yang, Zhaojun
Siohan, Olivier
Liu, Yue
Sun, Ming
Metze, Florian
Audio and Speech Processing
Sound
With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-computer interfaces. However, in real environments, interference from side-talk speech remains a significant challenge to WSR and may cause accumulated errors for downstream tasks such as natural language processing. In this work, we introduce a novel multi-channel differential automatic speech recognition (ASR) method for robust WSR on smart glasses. The proposed system takes differential inputs from different frontends that complement each other to improve the robustness of WSR, including a beamformer, microphone selection, and a lightweight side-talk detection model. Evaluations on both simulated and real datasets demonstrate that the proposed system outperforms the traditional approach, achieving up to an 18.0% relative reduction in word error rate.
title Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.14430