MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911188885962752 |
|---|---|
| author | Liu, Yinhong He, Jianfeng Su, Hang Lian, Ruixue Nian, Yi Vincent, Jake Vishnubhotla, Srikanth Piramuthu, Robinson Mansour, Saab |
| author_facet | Liu, Yinhong He, Jianfeng Su, Hang Lian, Ruixue Nian, Yi Vincent, Jake Vishnubhotla, Srikanth Piramuthu, Robinson Mansour, Saab |
| contents | Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark grounded in human annotations. In this work, we introduce MDSEval, the first meta-evaluation benchmark for MDS, consisting image-sharing dialogues, corresponding summaries, and human judgments across eight well-defined quality aspects. To ensure data quality and richfulness, we propose a novel filtering framework leveraging Mutually Exclusive Key Information (MEKI) across modalities. Our work is the first to identify and formalize key evaluation dimensions specific to MDS. We benchmark state-of-the-art modal evaluation methods, revealing their limitations in distinguishing summaries from advanced MLLMs and their susceptibility to various bias. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_01659 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization Liu, Yinhong He, Jianfeng Su, Hang Lian, Ruixue Nian, Yi Vincent, Jake Vishnubhotla, Srikanth Piramuthu, Robinson Mansour, Saab Computation and Language Artificial Intelligence Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark grounded in human annotations. In this work, we introduce MDSEval, the first meta-evaluation benchmark for MDS, consisting image-sharing dialogues, corresponding summaries, and human judgments across eight well-defined quality aspects. To ensure data quality and richfulness, we propose a novel filtering framework leveraging Mutually Exclusive Key Information (MEKI) across modalities. Our work is the first to identify and formalize key evaluation dimensions specific to MDS. We benchmark state-of-the-art modal evaluation methods, revealing their limitations in distinguishing summaries from advanced MLLMs and their susceptibility to various bias. |
| title | MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.01659 |