What's Wrong? Refining Meeting Summaries with LLM Feedback

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kirstein, Frederic, Ruas, Terry, Gipp, Bela
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910833528799232
author Kirstein, Frederic
Ruas, Terry
Gipp, Bela
author_facet Kirstein, Frederic
Ruas, Terry
Gipp, Bela
contents Meeting summarization has become a critical task since digital encounters have become a common practice. Large language models (LLMs) show great potential in summarization, offering enhanced coherence and context understanding compared to traditional methods. However, they still struggle to maintain relevance and avoid hallucination. We introduce a multi-LLM correction approach for meeting summarization using a two-phase process that mimics the human review process: mistake identification and summary refinement. We release QMSum Mistake, a dataset of 200 automatically generated meeting summaries annotated by humans on nine error types, including structural, omission, and irrelevance errors. Our experiments show that these errors can be identified with high accuracy by an LLM. We transform identified mistakes into actionable feedback to improve the quality of a given summary measured by relevance, informativeness, conciseness, and coherence. This post-hoc refinement effectively improves summary quality by leveraging multiple LLMs to validate output quality. Our multi-LLM approach for meeting summarization shows potential for similar complex text generation tasks requiring robustness, action planning, and discussion towards a goal.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11919
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle What's Wrong? Refining Meeting Summaries with LLM Feedback
Kirstein, Frederic
Ruas, Terry
Gipp, Bela
Computation and Language
Artificial Intelligence
Meeting summarization has become a critical task since digital encounters have become a common practice. Large language models (LLMs) show great potential in summarization, offering enhanced coherence and context understanding compared to traditional methods. However, they still struggle to maintain relevance and avoid hallucination. We introduce a multi-LLM correction approach for meeting summarization using a two-phase process that mimics the human review process: mistake identification and summary refinement. We release QMSum Mistake, a dataset of 200 automatically generated meeting summaries annotated by humans on nine error types, including structural, omission, and irrelevance errors. Our experiments show that these errors can be identified with high accuracy by an LLM. We transform identified mistakes into actionable feedback to improve the quality of a given summary measured by relevance, informativeness, conciseness, and coherence. This post-hoc refinement effectively improves summary quality by leveraging multiple LLMs to validate output quality. Our multi-LLM approach for meeting summarization shows potential for similar complex text generation tasks requiring robustness, action planning, and discussion towards a goal.
title What's Wrong? Refining Meeting Summaries with LLM Feedback
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.11919