An Empirical Study of Many-to-Many Summarization with Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Jiaan, Meng, Fandong, Sun, Zengkui, Liang, Yunlong, Cao, Yuxuan, Xu, Jiarong, Shi, Haoxiang, Zhou, Jie
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908370384977920
author Wang, Jiaan
Meng, Fandong
Sun, Zengkui
Liang, Yunlong
Cao, Yuxuan
Xu, Jiarong
Shi, Haoxiang
Zhou, Jie
author_facet Wang, Jiaan
Meng, Fandong
Sun, Zengkui
Liang, Yunlong
Cao, Yuxuan
Xu, Jiarong
Shi, Haoxiang
Zhou, Jie
contents Many-to-many summarization (M2MS) aims to process documents in any language and generate the corresponding summaries also in any language. Recently, large language models (LLMs) have shown strong multi-lingual abilities, giving them the potential to perform M2MS in real applications. This work presents a systematic empirical study on LLMs' M2MS ability. Specifically, we first reorganize M2MS data based on eight previous domain-specific datasets. The reorganized data contains 47.8K samples spanning five domains and six languages, which could be used to train and evaluate LLMs. Then, we benchmark 18 LLMs in a zero-shot manner and an instruction-tuning manner. Fine-tuned traditional models (e.g., mBART) are also conducted for comparisons. Our experiments reveal that, zero-shot LLMs achieve competitive results with fine-tuned traditional models. After instruct-tuning, open-source LLMs can significantly improve their M2MS ability, and outperform zero-shot LLMs (including GPT-4) in terms of automatic evaluations. In addition, we demonstrate that this task-specific improvement does not sacrifice the LLMs' general task-solving abilities. However, as revealed by our human evaluation, LLMs still face the factuality issue, and the instruction tuning might intensify the issue. Thus, how to control factual errors becomes the key when building LLM summarizers in real applications, and is worth noting in future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12983
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Study of Many-to-Many Summarization with Large Language Models
Wang, Jiaan
Meng, Fandong
Sun, Zengkui
Liang, Yunlong
Cao, Yuxuan
Xu, Jiarong
Shi, Haoxiang
Zhou, Jie
Computation and Language
Artificial Intelligence
Many-to-many summarization (M2MS) aims to process documents in any language and generate the corresponding summaries also in any language. Recently, large language models (LLMs) have shown strong multi-lingual abilities, giving them the potential to perform M2MS in real applications. This work presents a systematic empirical study on LLMs' M2MS ability. Specifically, we first reorganize M2MS data based on eight previous domain-specific datasets. The reorganized data contains 47.8K samples spanning five domains and six languages, which could be used to train and evaluate LLMs. Then, we benchmark 18 LLMs in a zero-shot manner and an instruction-tuning manner. Fine-tuned traditional models (e.g., mBART) are also conducted for comparisons. Our experiments reveal that, zero-shot LLMs achieve competitive results with fine-tuned traditional models. After instruct-tuning, open-source LLMs can significantly improve their M2MS ability, and outperform zero-shot LLMs (including GPT-4) in terms of automatic evaluations. In addition, we demonstrate that this task-specific improvement does not sacrifice the LLMs' general task-solving abilities. However, as revealed by our human evaluation, LLMs still face the factuality issue, and the instruction tuning might intensify the issue. Thus, how to control factual errors becomes the key when building LLM summarizers in real applications, and is worth noting in future research.
title An Empirical Study of Many-to-Many Summarization with Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.12983