An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aly, Walid Mohamed, Soliman, Taysir Hassan A., AbdelAziz, Amr Mohamed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909677582811136
author Aly, Walid Mohamed
Soliman, Taysir Hassan A.
AbdelAziz, Amr Mohamed
author_facet Aly, Walid Mohamed
Soliman, Taysir Hassan A.
AbdelAziz, Amr Mohamed
contents Large Language Models (LLMs) continue to advance natural language processing with their ability to generate human-like text across a range of tasks. Despite the remarkable success of LLMs in Natural Language Processing (NLP), their performance in text summarization across various domains and datasets has not been comprehensively evaluated. At the same time, the ability to summarize text effectively without relying on extensive training data has become a crucial bottleneck. To address these issues, we present a systematic evaluation of six LLMs across four datasets: CNN/Daily Mail and NewsRoom (news), SAMSum (dialog), and ArXiv (scientific). By leveraging prompt engineering techniques including zero-shot and in-context learning, our study evaluates the performance using the ROUGE and BERTScore metrics. In addition, a detailed analysis of inference times is conducted to better understand the trade-off between summarization quality and computational efficiency. For Long documents, introduce a sentence-based chunking strategy that enables LLMs with shorter context windows to summarize extended inputs in multiple stages. The findings reveal that while LLMs perform competitively on news and dialog tasks, their performance on long scientific documents improves significantly when aided by chunking strategies. In addition, notable performance variations were observed based on model parameters, dataset properties, and prompt design. These results offer actionable insights into how different LLMs behave across task types, contributing to ongoing research in efficient, instruction-based NLP systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05123
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
Aly, Walid Mohamed
Soliman, Taysir Hassan A.
AbdelAziz, Amr Mohamed
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) continue to advance natural language processing with their ability to generate human-like text across a range of tasks. Despite the remarkable success of LLMs in Natural Language Processing (NLP), their performance in text summarization across various domains and datasets has not been comprehensively evaluated. At the same time, the ability to summarize text effectively without relying on extensive training data has become a crucial bottleneck. To address these issues, we present a systematic evaluation of six LLMs across four datasets: CNN/Daily Mail and NewsRoom (news), SAMSum (dialog), and ArXiv (scientific). By leveraging prompt engineering techniques including zero-shot and in-context learning, our study evaluates the performance using the ROUGE and BERTScore metrics. In addition, a detailed analysis of inference times is conducted to better understand the trade-off between summarization quality and computational efficiency. For Long documents, introduce a sentence-based chunking strategy that enables LLMs with shorter context windows to summarize extended inputs in multiple stages. The findings reveal that while LLMs perform competitively on news and dialog tasks, their performance on long scientific documents improves significantly when aided by chunking strategies. In addition, notable performance variations were observed based on model parameters, dataset properties, and prompt design. These results offer actionable insights into how different LLMs behave across task types, contributing to ongoing research in efficient, instruction-based NLP systems.
title An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.05123