Proactive Hearing Assistants that Isolate Egocentric Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Guilin, Itani, Malek, Chen, Tuochao, Gollakota, Shyamnath
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908653584384000
author Hu, Guilin
Itani, Malek
Chen, Tuochao
Gollakota, Shyamnath
author_facet Hu, Guilin
Itani, Malek
Chen, Tuochao
Gollakota, Shyamnath
contents We introduce proactive hearing assistants that automatically identify and separate the wearer's conversation partners, without requiring explicit prompts. Our system operates on egocentric binaural audio and uses the wearer's self-speech as an anchor, leveraging turn-taking behavior and dialogue dynamics to infer conversational partners and suppress others. To enable real-time, on-device operation, we propose a dual-model architecture: a lightweight streaming model runs every 12.5 ms for low-latency extraction of the conversation partners, while a slower model runs less frequently to capture longer-range conversational dynamics. Results on real-world 2- and 3-speaker conversation test sets, collected with binaural egocentric hardware from 11 participants totaling 6.8 hours, show generalization in identifying and isolating conversational partners in multi-conversation settings. Our work marks a step toward hearing assistants that adapt proactively to conversational dynamics and engagement. More information can be found on our website: https://proactivehearing.cs.washington.edu/
format Preprint
id arxiv_https___arxiv_org_abs_2511_11473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Proactive Hearing Assistants that Isolate Egocentric Conversations
Hu, Guilin
Itani, Malek
Chen, Tuochao
Gollakota, Shyamnath
Computation and Language
Sound
Audio and Speech Processing
We introduce proactive hearing assistants that automatically identify and separate the wearer's conversation partners, without requiring explicit prompts. Our system operates on egocentric binaural audio and uses the wearer's self-speech as an anchor, leveraging turn-taking behavior and dialogue dynamics to infer conversational partners and suppress others. To enable real-time, on-device operation, we propose a dual-model architecture: a lightweight streaming model runs every 12.5 ms for low-latency extraction of the conversation partners, while a slower model runs less frequently to capture longer-range conversational dynamics. Results on real-world 2- and 3-speaker conversation test sets, collected with binaural egocentric hardware from 11 participants totaling 6.8 hours, show generalization in identifying and isolating conversational partners in multi-conversation settings. Our work marks a step toward hearing assistants that adapt proactively to conversational dynamics and engagement. More information can be found on our website: https://proactivehearing.cs.washington.edu/
title Proactive Hearing Assistants that Isolate Egocentric Conversations
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2511.11473