MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yinhong, He, Jianfeng, Su, Hang, Lian, Ruixue, Nian, Yi, Vincent, Jake, Vishnubhotla, Srikanth, Piramuthu, Robinson, Mansour, Saab
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911188885962752
author Liu, Yinhong
He, Jianfeng
Su, Hang
Lian, Ruixue
Nian, Yi
Vincent, Jake
Vishnubhotla, Srikanth
Piramuthu, Robinson
Mansour, Saab
author_facet Liu, Yinhong
He, Jianfeng
Su, Hang
Lian, Ruixue
Nian, Yi
Vincent, Jake
Vishnubhotla, Srikanth
Piramuthu, Robinson
Mansour, Saab
contents Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark grounded in human annotations. In this work, we introduce MDSEval, the first meta-evaluation benchmark for MDS, consisting image-sharing dialogues, corresponding summaries, and human judgments across eight well-defined quality aspects. To ensure data quality and richfulness, we propose a novel filtering framework leveraging Mutually Exclusive Key Information (MEKI) across modalities. Our work is the first to identify and formalize key evaluation dimensions specific to MDS. We benchmark state-of-the-art modal evaluation methods, revealing their limitations in distinguishing summaries from advanced MLLMs and their susceptibility to various bias.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01659
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
Liu, Yinhong
He, Jianfeng
Su, Hang
Lian, Ruixue
Nian, Yi
Vincent, Jake
Vishnubhotla, Srikanth
Piramuthu, Robinson
Mansour, Saab
Computation and Language
Artificial Intelligence
Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark grounded in human annotations. In this work, we introduce MDSEval, the first meta-evaluation benchmark for MDS, consisting image-sharing dialogues, corresponding summaries, and human judgments across eight well-defined quality aspects. To ensure data quality and richfulness, we propose a novel filtering framework leveraging Mutually Exclusive Key Information (MEKI) across modalities. Our work is the first to identify and formalize key evaluation dimensions specific to MDS. We benchmark state-of-the-art modal evaluation methods, revealing their limitations in distinguishing summaries from advanced MLLMs and their susceptibility to various bias.
title MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.01659