Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Chenghao, Luo, Yinbo, Wen, Zhoufutu, Chu, Qi, Gong, Tao, Liu, Longxiang, Zhang, Kaiyuan, Jiao, Jianpeng, Zhang, Ge, Huang, Wenhao, Yu, Nenghai
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2505.23810
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912585256796160
author Yang, Chenghao
Luo, Yinbo
Wen, Zhoufutu
Chu, Qi
Gong, Tao
Liu, Longxiang
Zhang, Kaiyuan
Jiao, Jianpeng
Zhang, Ge
Huang, Wenhao
Yu, Nenghai
author_facet Yang, Chenghao
Luo, Yinbo
Wen, Zhoufutu
Chu, Qi
Gong, Tao
Liu, Longxiang
Zhang, Kaiyuan
Jiao, Jianpeng
Zhang, Ge
Huang, Wenhao
Yu, Nenghai
contents Large Language Models (\textbf{LLMs}), e.g. ChatGPT, have been widely adopted in real-world dialogue applications. However, LLMs' robustness, especially in handling long complex dialogue sessions, including frequent motivation transfer, sophisticated cross-turn dependency, is criticized all along. Nevertheless, no existing benchmarks can fully reflect these weaknesses. We present \textbf{MARS-Bench}, a \textbf{M}ulti-turn \textbf{A}thletic \textbf{R}eal-world \textbf{S}cenario Dialogue \textbf{Bench}mark, designed to remedy the gap. MARS-Bench is constructed from play-by-play text commentary so to feature realistic dialogues specifically designed to evaluate three critical aspects of multi-turn conversations: Ultra Multi-turn, Interactive Multi-turn, and Cross-turn Tasks. Extensive experiments on MARS-Bench also reveal that closed-source LLMs significantly outperform open-source alternatives, explicit reasoning significantly boosts LLMs' robustness on handling long complex dialogue sessions, and LLMs indeed face significant challenges when handling motivation transfer and sophisticated cross-turn dependency. Moreover, we provide mechanistic interpretability on how attention sinks due to special tokens lead to LLMs' performance degradation when handling long complex dialogue sessions based on attention visualization experiment in Qwen2.5-7B-Instruction.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
Yang, Chenghao
Luo, Yinbo
Wen, Zhoufutu
Chu, Qi
Gong, Tao
Liu, Longxiang
Zhang, Kaiyuan
Jiao, Jianpeng
Zhang, Ge
Huang, Wenhao
Yu, Nenghai
Computation and Language
Artificial Intelligence
Large Language Models (\textbf{LLMs}), e.g. ChatGPT, have been widely adopted in real-world dialogue applications. However, LLMs' robustness, especially in handling long complex dialogue sessions, including frequent motivation transfer, sophisticated cross-turn dependency, is criticized all along. Nevertheless, no existing benchmarks can fully reflect these weaknesses. We present \textbf{MARS-Bench}, a \textbf{M}ulti-turn \textbf{A}thletic \textbf{R}eal-world \textbf{S}cenario Dialogue \textbf{Bench}mark, designed to remedy the gap. MARS-Bench is constructed from play-by-play text commentary so to feature realistic dialogues specifically designed to evaluate three critical aspects of multi-turn conversations: Ultra Multi-turn, Interactive Multi-turn, and Cross-turn Tasks. Extensive experiments on MARS-Bench also reveal that closed-source LLMs significantly outperform open-source alternatives, explicit reasoning significantly boosts LLMs' robustness on handling long complex dialogue sessions, and LLMs indeed face significant challenges when handling motivation transfer and sophisticated cross-turn dependency. Moreover, we provide mechanistic interpretability on how attention sinks due to special tokens lead to LLMs' performance degradation when handling long complex dialogue sessions based on attention visualization experiment in Qwen2.5-7B-Instruction.
title MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.23810