Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fonseca, Marcio, Cohen, Shay B.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914849601093632
author Fonseca, Marcio
Cohen, Shay B.
author_facet Fonseca, Marcio
Cohen, Shay B.
contents In this work, we investigate the controllability of large language models (LLMs) on scientific summarization tasks. We identify key stylistic and content coverage factors that characterize different types of summaries such as paper reviews, abstracts, and lay summaries. By controlling stylistic features, we find that non-fine-tuned LLMs outperform humans in the MuP review generation task, both in terms of similarity to reference summaries and human preferences. Also, we show that we can improve the controllability of LLMs with keyword-based classifier-free guidance (CFG) while achieving lexical overlap comparable to strong fine-tuned baselines on arXiv and PubMed. However, our results also indicate that LLMs cannot consistently generate long summaries with more than 8 sentences. Furthermore, these models exhibit limited capacity to produce highly abstractive lay summaries. Although LLMs demonstrate strong generic summarization competency, sophisticated content control without costly fine-tuning remains an open problem for domain-specific applications.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10415
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
Fonseca, Marcio
Cohen, Shay B.
Computation and Language
Artificial Intelligence
In this work, we investigate the controllability of large language models (LLMs) on scientific summarization tasks. We identify key stylistic and content coverage factors that characterize different types of summaries such as paper reviews, abstracts, and lay summaries. By controlling stylistic features, we find that non-fine-tuned LLMs outperform humans in the MuP review generation task, both in terms of similarity to reference summaries and human preferences. Also, we show that we can improve the controllability of LLMs with keyword-based classifier-free guidance (CFG) while achieving lexical overlap comparable to strong fine-tuned baselines on arXiv and PubMed. However, our results also indicate that LLMs cannot consistently generate long summaries with more than 8 sentences. Furthermore, these models exhibit limited capacity to produce highly abstractive lay summaries. Although LLMs demonstrate strong generic summarization competency, sophisticated content control without costly fine-tuning remains an open problem for domain-specific applications.
title Can Large Language Model Summarizers Adapt to Diverse Scientific Communication Goals?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.10415