SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Kuang, Wang, Yifeng, Zhang, Xiyuxing, Shen, Chengyi, Kumar, Swarun, Chan, Justin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912898228420608
author Yuan, Kuang
Wang, Yifeng
Zhang, Xiyuxing
Shen, Chengyi
Kumar, Swarun
Chan, Justin
author_facet Yuan, Kuang
Wang, Yifeng
Zhang, Xiyuxing
Shen, Chengyi
Kumar, Swarun
Chan, Justin
contents Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30° angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
Yuan, Kuang
Wang, Yifeng
Zhang, Xiyuxing
Shen, Chengyi
Kumar, Swarun
Chan, Justin
Sound
Human-Computer Interaction
Machine Learning
Audio and Speech Processing
Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30° angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays.
title SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
topic Sound
Human-Computer Interaction
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2504.10793