Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911160793563136 |
|---|---|
| author | Yang, Yufeng Huang, Yiteng Xu, Yong Wan, Li Shon, Suwon Liu, Yang Fan, Yifeng Yang, Zhaojun Siohan, Olivier Liu, Yue Sun, Ming Metze, Florian |
| author_facet | Yang, Yufeng Huang, Yiteng Xu, Yong Wan, Li Shon, Suwon Liu, Yang Fan, Yifeng Yang, Zhaojun Siohan, Olivier Liu, Yue Sun, Ming Metze, Florian |
| contents | With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-computer interfaces. However, in real environments, interference from side-talk speech remains a significant challenge to WSR and may cause accumulated errors for downstream tasks such as natural language processing. In this work, we introduce a novel multi-channel differential automatic speech recognition (ASR) method for robust WSR on smart glasses. The proposed system takes differential inputs from different frontends that complement each other to improve the robustness of WSR, including a beamformer, microphone selection, and a lightweight side-talk detection model. Evaluations on both simulated and real datasets demonstrate that the proposed system outperforms the traditional approach, achieving up to an 18.0% relative reduction in word error rate. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14430 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses Yang, Yufeng Huang, Yiteng Xu, Yong Wan, Li Shon, Suwon Liu, Yang Fan, Yifeng Yang, Zhaojun Siohan, Olivier Liu, Yue Sun, Ming Metze, Florian Audio and Speech Processing Sound With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-computer interfaces. However, in real environments, interference from side-talk speech remains a significant challenge to WSR and may cause accumulated errors for downstream tasks such as natural language processing. In this work, we introduce a novel multi-channel differential automatic speech recognition (ASR) method for robust WSR on smart glasses. The proposed system takes differential inputs from different frontends that complement each other to improve the robustness of WSR, including a beamformer, microphone selection, and a lightweight side-talk detection model. Evaluations on both simulated and real datasets demonstrate that the proposed system outperforms the traditional approach, achieving up to an 18.0% relative reduction in word error rate. |
| title | Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2509.14430 |