WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912803952001024 |
|---|---|
| author | Lin, Zhaojiang Xu, Yong Sun, Kai Zheng, Jing Huang, Yin Appini, Surya Teja Narang, Krish Tao, Renjie Jain, Ishan Kapil Arora, Siddhant Li, Ruizhi Huang, Yiteng Patnaik, Kaushik Xu, Wenfang Shon, Suwon Liu, Yue Aly, Ahmed A Kumar, Anuj Metze, Florian Dong, Xin Luna |
| author_facet | Lin, Zhaojiang Xu, Yong Sun, Kai Zheng, Jing Huang, Yin Appini, Surya Teja Narang, Krish Tao, Renjie Jain, Ishan Kapil Arora, Siddhant Li, Ruizhi Huang, Yiteng Patnaik, Kaushik Xu, Wenfang Shon, Suwon Liu, Yue Aly, Ahmed A Kumar, Anuj Metze, Florian Dong, Xin Luna |
| contents | Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_02391 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables Lin, Zhaojiang Xu, Yong Sun, Kai Zheng, Jing Huang, Yin Appini, Surya Teja Narang, Krish Tao, Renjie Jain, Ishan Kapil Arora, Siddhant Li, Ruizhi Huang, Yiteng Patnaik, Kaushik Xu, Wenfang Shon, Suwon Liu, Yue Aly, Ahmed A Kumar, Anuj Metze, Florian Dong, Xin Luna Computation and Language Sound Audio and Speech Processing Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research. |
| title | WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2601.02391 |