TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zehan, Chen, Hongjie, Wang, Qing, Zhang, Yuxin, Zhou, Jing, Lv, Hang, Du, Mengjie, Song, Yaodong, Lian, Jie, Kang, Jian, Li, Jie, Li, Yongxiang, Li, Xuelong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909987137126400
author Li, Zehan
Chen, Hongjie
Wang, Qing
Zhang, Yuxin
Zhou, Jing
Lv, Hang
Du, Mengjie
Song, Yaodong
Lian, Jie
Kang, Jian
Li, Jie
Li, Yongxiang
Li, Xuelong
author_facet Li, Zehan
Chen, Hongjie
Wang, Qing
Zhang, Yuxin
Zhou, Jing
Lv, Hang
Du, Mengjie
Song, Yaodong
Lian, Jie
Kang, Jian
Li, Jie
Li, Yongxiang
Li, Xuelong
contents Spoken language models (SLMs) have advanced rapidly in recent years, accompanied by a growing number of evaluation benchmarks. However, most existing benchmarks emphasize task completion and capability scaling, while remaining poorly aligned with how users interact with SLMs in real-world spoken conversations. Effective spoken interaction requires not only accurate understanding of user intent and content, but also the ability to respond with appropriate interactional strategies. In this paper, we present TELEVAL, a dynamic, user-centered benchmark for evaluating SLMs in realistic Chinese spoken interaction scenarios. TELEVAL consolidates evaluation into two core aspects. Reliable Content Fulfillment assesses whether models can comprehend spoken inputs and produce semantically correct responses. Interactional Appropriateness evaluates whether models act as socially capable interlocutors, requiring them not only to generate human-like, colloquial responses, but also to implicitly incorporate paralinguistic cues for natural interaction. Experiments reveal that, despite strong performance on semantic and knowledge-oriented tasks, current SLMs still struggle to produce natural and interactionally appropriate responses, highlighting the need for more interaction-faithful evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18061
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
Li, Zehan
Chen, Hongjie
Wang, Qing
Zhang, Yuxin
Zhou, Jing
Lv, Hang
Du, Mengjie
Song, Yaodong
Lian, Jie
Kang, Jian
Li, Jie
Li, Yongxiang
Li, Xuelong
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Spoken language models (SLMs) have advanced rapidly in recent years, accompanied by a growing number of evaluation benchmarks. However, most existing benchmarks emphasize task completion and capability scaling, while remaining poorly aligned with how users interact with SLMs in real-world spoken conversations. Effective spoken interaction requires not only accurate understanding of user intent and content, but also the ability to respond with appropriate interactional strategies. In this paper, we present TELEVAL, a dynamic, user-centered benchmark for evaluating SLMs in realistic Chinese spoken interaction scenarios. TELEVAL consolidates evaluation into two core aspects. Reliable Content Fulfillment assesses whether models can comprehend spoken inputs and produce semantically correct responses. Interactional Appropriateness evaluates whether models act as socially capable interlocutors, requiring them not only to generate human-like, colloquial responses, but also to implicitly incorporate paralinguistic cues for natural interaction. Experiments reveal that, despite strong performance on semantic and knowledge-oriented tasks, current SLMs still struggle to produce natural and interactionally appropriate responses, highlighting the need for more interaction-faithful evaluation.
title TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2507.18061