Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Muppidi, Ananth, Das, Tarak, Bandyopadhyay, Sambaran, Shukla, Tripti, A, Dharun D
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908378210500608
author Muppidi, Ananth
Das, Tarak
Bandyopadhyay, Sambaran
Shukla, Tripti
A, Dharun D
author_facet Muppidi, Ananth
Das, Tarak
Bandyopadhyay, Sambaran
Shukla, Tripti
A, Dharun D
contents The generation of presentation slides automatically is an important problem in the era of generative AI. This paper focuses on evaluating multimodal content in presentation slides that can effectively summarize a document and convey concepts to a broad audience. We introduce a benchmark dataset, RefSlides, consisting of human-made high-quality presentations that span various topics. Next, we propose a set of metrics to characterize different intrinsic properties of the content of a presentation and present REFLEX, an evaluation approach that generates scores and actionable feedback for these metrics. We achieve this by generating negative presentation samples with different degrees of metric-specific perturbations and use them to fine-tune LLMs. This reference-free evaluation technique does not require ground truth presentations during inference. Our extensive automated and human experiments demonstrate that our evaluation approach outperforms classical heuristic-based and state-of-the-art large language model-based evaluations in generating scores and explanations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18240
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
Muppidi, Ananth
Das, Tarak
Bandyopadhyay, Sambaran
Shukla, Tripti
A, Dharun D
Computation and Language
Artificial Intelligence
The generation of presentation slides automatically is an important problem in the era of generative AI. This paper focuses on evaluating multimodal content in presentation slides that can effectively summarize a document and convey concepts to a broad audience. We introduce a benchmark dataset, RefSlides, consisting of human-made high-quality presentations that span various topics. Next, we propose a set of metrics to characterize different intrinsic properties of the content of a presentation and present REFLEX, an evaluation approach that generates scores and actionable feedback for these metrics. We achieve this by generating negative presentation samples with different degrees of metric-specific perturbations and use them to fine-tune LLMs. This reference-free evaluation technique does not require ground truth presentations during inference. Our extensive automated and human experiments demonstrate that our evaluation approach outperforms classical heuristic-based and state-of-the-art large language model-based evaluations in generating scores and explanations.
title Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.18240