Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rallapalli, Swati, Gallagher, Shannon, Mellinger, Andrew O., Ratchford, Jasmine, Sinha, Anusha, Brooks, Tyler, Nichols, William R., Winski, Nick, Brown, Bryan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908960753188864
author Rallapalli, Swati
Gallagher, Shannon
Mellinger, Andrew O.
Ratchford, Jasmine
Sinha, Anusha
Brooks, Tyler
Nichols, William R.
Winski, Nick
Brown, Bryan
author_facet Rallapalli, Swati
Gallagher, Shannon
Mellinger, Andrew O.
Ratchford, Jasmine
Sinha, Anusha
Brooks, Tyler
Nichols, William R.
Winski, Nick
Brown, Bryan
contents We study the efficacy of fine-tuning Large Language Models (LLMs) for the specific task of report (government archives, news, intelligence reports) summarization. While this topic is being very actively researched - our specific application set-up faces two challenges: (i) ground-truth summaries maybe unavailable (e.g., for government archives), and (ii) availability of limited compute power - the sensitive nature of the application requires that computation is performed on-premise and for most of our experiments we use one or two A100 GPU cards. Under this set-up we conduct experiments to answer the following questions. First, given that fine-tuning the LLMs can be resource intensive, is it feasible to fine-tune them for improved report summarization capabilities on-premise? Second, what are the metrics we could leverage to assess the quality of these summaries? We conduct experiments on two different fine-tuning approaches in parallel and our findings reveal interesting trends regarding the utility of fine-tuning LLMs. Specifically, we find that in many cases, fine-tuning helps improve summary quality and in other cases it helps by reducing the number of invalid or garbage summaries.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data
Rallapalli, Swati
Gallagher, Shannon
Mellinger, Andrew O.
Ratchford, Jasmine
Sinha, Anusha
Brooks, Tyler
Nichols, William R.
Winski, Nick
Brown, Bryan
Computation and Language
Artificial Intelligence
Machine Learning
We study the efficacy of fine-tuning Large Language Models (LLMs) for the specific task of report (government archives, news, intelligence reports) summarization. While this topic is being very actively researched - our specific application set-up faces two challenges: (i) ground-truth summaries maybe unavailable (e.g., for government archives), and (ii) availability of limited compute power - the sensitive nature of the application requires that computation is performed on-premise and for most of our experiments we use one or two A100 GPU cards. Under this set-up we conduct experiments to answer the following questions. First, given that fine-tuning the LLMs can be resource intensive, is it feasible to fine-tune them for improved report summarization capabilities on-premise? Second, what are the metrics we could leverage to assess the quality of these summaries? We conduct experiments on two different fine-tuning approaches in parallel and our findings reveal interesting trends regarding the utility of fine-tuning LLMs. Specifically, we find that in many cases, fine-tuning helps improve summary quality and in other cases it helps by reducing the number of invalid or garbage summaries.
title Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.10676