DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ramnath, Sahana, Chitsazan, Nima, Zhou, Mingyang, Lee, Chia-Hsuan, Zhang, Shi-Xiong, Rawls, Stephen, Sahu, Sambit, Cho, Sangwoo, Ren, Xiang, Winata, Genta Indra, Veldanda, Akshaj Kumar
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917258081599488
author Ramnath, Sahana
Chitsazan, Nima
Zhou, Mingyang
Lee, Chia-Hsuan
Zhang, Shi-Xiong
Rawls, Stephen
Sahu, Sambit
Cho, Sangwoo
Ren, Xiang
Winata, Genta Indra
Veldanda, Akshaj Kumar
author_facet Ramnath, Sahana
Chitsazan, Nima
Zhou, Mingyang
Lee, Chia-Hsuan
Zhang, Shi-Xiong
Rawls, Stephen
Sahu, Sambit
Cho, Sangwoo
Ren, Xiang
Winata, Genta Indra
Veldanda, Akshaj Kumar
contents Dialogues are a predominant mode of communication for humans, and it is immensely helpful to have automatically generated summaries of them (e.g., to revise key points discussed in a meeting, to review conversations between customer agents and product users). Prior works on dialogue summary evaluation largely ignore the complexities specific to this task: (i) shift in structure, from multiple speakers discussing information in a scattered fashion across several turns, to a summary's sentences, and (ii) shift in narration viewpoint, from speakers' first/second-person narration, standardized third-person narration in the summary. In this work, we introduce our framework DIALSUMMER to address the above. We propose DIAL-SUMMER's taxonomy of errors to comprehensively evaluate dialogue summaries at two hierarchical levels: DIALOGUE-LEVEL that focuses on the broader speakers/turns, and WITHIN-TURN-LEVEL that focuses on the information talked about inside a turn. We then present DIAL-SUMMER's dataset composed of dialogue summaries manually annotated with our taxonomy's fine-grained errors. We conduct empirical analyses of these annotated errors, and observe interesting trends (e.g., turns occurring in middle of the dialogue are the most frequently missed in the summary, extrinsic hallucinations largely occur at the end of the summary). We also conduct experiments on LLM-Judges' capability at detecting these errors, through which we demonstrate the challenging nature of our dataset, the robustness of our taxonomy, and the need for future work in this field to enhance LLMs' performance in the same. Code and inference dataset coming soon.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08149
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
Ramnath, Sahana
Chitsazan, Nima
Zhou, Mingyang
Lee, Chia-Hsuan
Zhang, Shi-Xiong
Rawls, Stephen
Sahu, Sambit
Cho, Sangwoo
Ren, Xiang
Winata, Genta Indra
Veldanda, Akshaj Kumar
Computation and Language
Artificial Intelligence
Dialogues are a predominant mode of communication for humans, and it is immensely helpful to have automatically generated summaries of them (e.g., to revise key points discussed in a meeting, to review conversations between customer agents and product users). Prior works on dialogue summary evaluation largely ignore the complexities specific to this task: (i) shift in structure, from multiple speakers discussing information in a scattered fashion across several turns, to a summary's sentences, and (ii) shift in narration viewpoint, from speakers' first/second-person narration, standardized third-person narration in the summary. In this work, we introduce our framework DIALSUMMER to address the above. We propose DIAL-SUMMER's taxonomy of errors to comprehensively evaluate dialogue summaries at two hierarchical levels: DIALOGUE-LEVEL that focuses on the broader speakers/turns, and WITHIN-TURN-LEVEL that focuses on the information talked about inside a turn. We then present DIAL-SUMMER's dataset composed of dialogue summaries manually annotated with our taxonomy's fine-grained errors. We conduct empirical analyses of these annotated errors, and observe interesting trends (e.g., turns occurring in middle of the dialogue are the most frequently missed in the summary, extrinsic hallucinations largely occur at the end of the summary). We also conduct experiments on LLM-Judges' capability at detecting these errors, through which we demonstrate the challenging nature of our dataset, the robustness of our taxonomy, and the need for future work in this field to enhance LLMs' performance in the same. Code and inference dataset coming soon.
title DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.08149