LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, Yongxuan, Chen, Runyu, Liu, Peiyu, Qian, Hongjin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908336224468992
author Wu, Yongxuan
Chen, Runyu
Liu, Peiyu
Qian, Hongjin
author_facet Wu, Yongxuan
Chen, Runyu
Liu, Peiyu
Qian, Hongjin
contents Long-context understanding poses significant challenges in natural language processing, particularly for real-world dialogues characterized by speech-based elements, high redundancy, and uneven information density. Although large language models (LLMs) achieve impressive results on existing benchmarks, these datasets fail to reflect the complexities of such texts, limiting their applicability to practical scenarios. To bridge this gap, we construct the first spoken long-text dataset, derived from live streams, designed to reflect the redundancy-rich and conversational nature of real-world scenarios. We construct tasks in three categories: retrieval-dependent, reasoning-dependent, and hybrid. We then evaluate both popular LLMs and specialized methods to assess their ability to understand long-contexts in these tasks. Our results show that current methods exhibit strong task-specific preferences and perform poorly on highly redundant inputs, with no single method consistently outperforming others. We propose a new baseline that better handles redundancy in spoken text and achieves strong performance across tasks. Our findings highlight key limitations of current methods and suggest future directions for improving long-context understanding. Finally, our benchmark fills a gap in evaluating long-context spoken language understanding and provides a practical foundation for developing real-world e-commerce systems. The code and benchmark are available at https://github.com/Yarayx/livelongbench.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17366
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams
Wu, Yongxuan
Chen, Runyu
Liu, Peiyu
Qian, Hongjin
Computation and Language
Artificial Intelligence
Long-context understanding poses significant challenges in natural language processing, particularly for real-world dialogues characterized by speech-based elements, high redundancy, and uneven information density. Although large language models (LLMs) achieve impressive results on existing benchmarks, these datasets fail to reflect the complexities of such texts, limiting their applicability to practical scenarios. To bridge this gap, we construct the first spoken long-text dataset, derived from live streams, designed to reflect the redundancy-rich and conversational nature of real-world scenarios. We construct tasks in three categories: retrieval-dependent, reasoning-dependent, and hybrid. We then evaluate both popular LLMs and specialized methods to assess their ability to understand long-contexts in these tasks. Our results show that current methods exhibit strong task-specific preferences and perform poorly on highly redundant inputs, with no single method consistently outperforming others. We propose a new baseline that better handles redundancy in spoken text and achieves strong performance across tasks. Our findings highlight key limitations of current methods and suggest future directions for improving long-context understanding. Finally, our benchmark fills a gap in evaluating long-context spoken language understanding and provides a practical foundation for developing real-world e-commerce systems. The code and benchmark are available at https://github.com/Yarayx/livelongbench.
title LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.17366