WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Zhaojiang, Xu, Yong, Sun, Kai, Zheng, Jing, Huang, Yin, Appini, Surya Teja, Narang, Krish, Tao, Renjie, Jain, Ishan Kapil, Arora, Siddhant, Li, Ruizhi, Huang, Yiteng, Patnaik, Kaushik, Xu, Wenfang, Shon, Suwon, Liu, Yue, Aly, Ahmed A, Kumar, Anuj, Metze, Florian, Dong, Xin Luna
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912803952001024
author Lin, Zhaojiang
Xu, Yong
Sun, Kai
Zheng, Jing
Huang, Yin
Appini, Surya Teja
Narang, Krish
Tao, Renjie
Jain, Ishan Kapil
Arora, Siddhant
Li, Ruizhi
Huang, Yiteng
Patnaik, Kaushik
Xu, Wenfang
Shon, Suwon
Liu, Yue
Aly, Ahmed A
Kumar, Anuj
Metze, Florian
Dong, Xin Luna
author_facet Lin, Zhaojiang
Xu, Yong
Sun, Kai
Zheng, Jing
Huang, Yin
Appini, Surya Teja
Narang, Krish
Tao, Renjie
Jain, Ishan Kapil
Arora, Siddhant
Li, Ruizhi
Huang, Yiteng
Patnaik, Kaushik
Xu, Wenfang
Shon, Suwon
Liu, Yue
Aly, Ahmed A
Kumar, Anuj
Metze, Florian
Dong, Xin Luna
contents Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
Lin, Zhaojiang
Xu, Yong
Sun, Kai
Zheng, Jing
Huang, Yin
Appini, Surya Teja
Narang, Krish
Tao, Renjie
Jain, Ishan Kapil
Arora, Siddhant
Li, Ruizhi
Huang, Yiteng
Patnaik, Kaushik
Xu, Wenfang
Shon, Suwon
Liu, Yue
Aly, Ahmed A
Kumar, Anuj
Metze, Florian
Dong, Xin Luna
Computation and Language
Sound
Audio and Speech Processing
Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.
title WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.02391