MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yinhong, He, Jianfeng, Su, Hang, Lian, Ruixue, Nian, Yi, Vincent, Jake, Vishnubhotla, Srikanth, Piramuthu, Robinson, Mansour, Saab |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024)
by: Tang, Liyan, et al.
Published: (2024)
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024)
by: Song, Hwanjun, et al.
Published: (2024)
GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback
by: Zou, Henry Peng, et al.
Published: (2025)
by: Zou, Henry Peng, et al.
Published: (2025)
Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
by: Koupaee, Mahnaz, et al.
Published: (2025)
by: Koupaee, Mahnaz, et al.
Published: (2025)
Controllable Conversational Theme Detection Track at DSTC 12
by: Shalyminov, Igor, et al.
Published: (2025)
by: Shalyminov, Igor, et al.
Published: (2025)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate
by: Yao, Binwei, et al.
Published: (2025)
by: Yao, Binwei, et al.
Published: (2025)
DFlow: Diverse Dialogue Flow Simulation with Large Language Models
by: Du, Wanyu, et al.
Published: (2024)
by: Du, Wanyu, et al.
Published: (2024)
CERET: Cost-Effective Extrinsic Refinement for Text Generation
by: Cai, Jason, et al.
Published: (2024)
by: Cai, Jason, et al.
Published: (2024)
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
by: Sheth, Ivaxi, et al.
Published: (2026)
by: Sheth, Ivaxi, et al.
Published: (2026)
Can Your Model Tell a Negation from an Implicature? Unravelling Challenges With Intent Encoders
by: Zhang, Yuwei, et al.
Published: (2024)
by: Zhang, Yuwei, et al.
Published: (2024)
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
PSentScore: Evaluating Sentiment Polarity in Dialogue Summarization
by: Zhou, Yongxin, et al.
Published: (2023)
by: Zhou, Yongxin, et al.
Published: (2023)
RADIUS: Ranking, Distribution, and Significance - A Comprehensive Alignment Suite for Survey Simulation
by: Łajewska, Weronika, et al.
Published: (2026)
by: Łajewska, Weronika, et al.
Published: (2026)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
by: Omar, Reham, et al.
Published: (2025)
by: Omar, Reham, et al.
Published: (2025)
Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings
by: Alqahtani, Sawsan, et al.
Published: (2021)
by: Alqahtani, Sawsan, et al.
Published: (2021)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
by: Nejadgholi, Isar, et al.
Published: (2026)
by: Nejadgholi, Isar, et al.
Published: (2026)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
by: Zhang, Junyi, et al.
Published: (2025)
by: Zhang, Junyi, et al.
Published: (2025)
Instructive Dialogue Summarization with Query Aggregations
by: Wang, Bin, et al.
Published: (2023)
by: Wang, Bin, et al.
Published: (2023)
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization
by: Lu, Yen-Ju, et al.
Published: (2025)
by: Lu, Yen-Ju, et al.
Published: (2025)
QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
by: Ghebriout, Mohamed Imed Eddine, et al.
Published: (2025)
by: Ghebriout, Mohamed Imed Eddine, et al.
Published: (2025)
Structured List-Grounded Question Answering
by: Sung, Mujeen, et al.
Published: (2024)
by: Sung, Mujeen, et al.
Published: (2024)
Multilingual Self-Taught Faithfulness Evaluators
by: Alfano, Carlo, et al.
Published: (2025)
by: Alfano, Carlo, et al.
Published: (2025)
FLAP: Flow-Adhering Planning with Constrained Decoding in LLMs
by: Roy, Shamik, et al.
Published: (2024)
by: Roy, Shamik, et al.
Published: (2024)
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
by: Suresh, Sathya Krishnan, et al.
Published: (2025)
by: Suresh, Sathya Krishnan, et al.
Published: (2025)
UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs
by: Lee, Yuho, et al.
Published: (2024)
by: Lee, Yuho, et al.
Published: (2024)
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation
by: Mendonça, John, et al.
Published: (2024)
by: Mendonça, John, et al.
Published: (2024)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
by: Park, Seongmin, et al.
Published: (2024)
by: Park, Seongmin, et al.
Published: (2024)
MAGID: An Automated Pipeline for Generating Synthetic Multi-modal Datasets
by: Aboutalebi, Hossein, et al.
Published: (2024)
by: Aboutalebi, Hossein, et al.
Published: (2024)
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories
by: Huang, Xi Yu, et al.
Published: (2024)
by: Huang, Xi Yu, et al.
Published: (2024)
Prompt Compression for Large Language Models: A Survey
by: Li, Zongqian, et al.
Published: (2024)
by: Li, Zongqian, et al.
Published: (2024)
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
Mitigating Semantic Drift: Evaluating LLMs' Efficacy in Psychotherapy through MI Dialogue Summarization
by: Kumar, Vivek, et al.
Published: (2025)
by: Kumar, Vivek, et al.
Published: (2025)
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
by: Chen, Kedi, et al.
Published: (2024)
by: Chen, Kedi, et al.
Published: (2024)
Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
by: He, Jianfeng, et al.
Published: (2024)
by: He, Jianfeng, et al.
Published: (2024)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
by: Peri, Raghuveer, et al.
Published: (2024)
by: Peri, Raghuveer, et al.
Published: (2024)
Eliciting Better Multilingual Structured Reasoning from LLMs through Code
by: Li, Bryan, et al.
Published: (2024)
by: Li, Bryan, et al.
Published: (2024)
SPECTRUM: Speaker-Enhanced Pre-Training for Long Dialogue Summarization
by: Cho, Sangwoo, et al.
Published: (2024)
by: Cho, Sangwoo, et al.
Published: (2024)
Similar Items
-
Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
by: He, Jianfeng, et al.
Published: (2024) -
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
by: Tang, Liyan, et al.
Published: (2024) -
FineSurE: Fine-grained Summarization Evaluation using LLMs
by: Song, Hwanjun, et al.
Published: (2024) -
GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback
by: Zou, Henry Peng, et al.
Published: (2025) -
Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
by: Koupaee, Mahnaz, et al.
Published: (2025)