SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912898228420608 |
|---|---|
| author | Yuan, Kuang Wang, Yifeng Zhang, Xiyuxing Shen, Chengyi Kumar, Swarun Chan, Justin |
| author_facet | Yuan, Kuang Wang, Yifeng Zhang, Xiyuxing Shen, Chengyi Kumar, Swarun Chan, Justin |
| contents | Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30° angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_10793 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures Yuan, Kuang Wang, Yifeng Zhang, Xiyuxing Shen, Chengyi Kumar, Swarun Chan, Justin Sound Human-Computer Interaction Machine Learning Audio and Speech Processing Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30° angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays. |
| title | SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures |
| topic | Sound Human-Computer Interaction Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2504.10793 |