Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Guan-Ting, Kuan, Shih-Yun Shan, Wang, Qirui, Lian, Jiachen, Li, Tingle, Watanabe, Shinji, Lee, Hung-yi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913060497653760
author Lin, Guan-Ting
Kuan, Shih-Yun Shan
Wang, Qirui
Lian, Jiachen
Li, Tingle
Watanabe, Shinji
Lee, Hung-yi
author_facet Lin, Guan-Ting
Kuan, Shih-Yun Shan
Wang, Qirui
Lian, Jiachen
Li, Tingle
Watanabe, Shinji
Lee, Hung-yi
contents Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech, remains critically under-evaluated. We introduce Full-Duplex-Bench v1.5, the first fully automated benchmark designed to systematically probe how models behave during speech overlap. The benchmark simulates four representative overlap scenarios: user interruption, user backchannel, talking to others, and background speech. Our framework, compatible with open-source and commercial API-based models, provides a comprehensive suite of metrics analyzing categorical dialogue behaviors, stop and response latency, and prosodic adaptation. Benchmarking five state-of-the-art agents reveals two divergent strategies: a responsive approach prioritizing rapid response to user input, and a floor-holding approach that preserves conversational flow by filtering overlapping events. Our open-source framework enables practitioners to accelerate the development of robust full-duplex systems by providing the tools for reproducible evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23159
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
Lin, Guan-Ting
Kuan, Shih-Yun Shan
Wang, Qirui
Lian, Jiachen
Li, Tingle
Watanabe, Shinji
Lee, Hung-yi
Audio and Speech Processing
Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech, remains critically under-evaluated. We introduce Full-Duplex-Bench v1.5, the first fully automated benchmark designed to systematically probe how models behave during speech overlap. The benchmark simulates four representative overlap scenarios: user interruption, user backchannel, talking to others, and background speech. Our framework, compatible with open-source and commercial API-based models, provides a comprehensive suite of metrics analyzing categorical dialogue behaviors, stop and response latency, and prosodic adaptation. Benchmarking five state-of-the-art agents reveals two divergent strategies: a responsive approach prioritizing rapid response to user input, and a floor-holding approach that preserves conversational flow by filtering overlapping events. Our open-source framework enables practitioners to accelerate the development of robust full-duplex systems by providing the tools for reproducible evaluation.
title Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
topic Audio and Speech Processing
url https://arxiv.org/abs/2507.23159