Comparing Approaches to Automatic Summarization in Less-Resourced Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Palen-Michel, Chester, Lignos, Constantine
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909978841841664
author Palen-Michel, Chester
Lignos, Constantine
author_facet Palen-Michel, Chester
Lignos, Constantine
contents Automatic text summarization has achieved high performance in high-resourced languages like English, but comparatively less attention has been given to summarization in less-resourced languages. This work compares a variety of different approaches to summarization from zero-shot prompting of LLMs large and small to fine-tuning smaller models like mT5 with and without three data augmentation approaches and multilingual transfer. We also explore an LLM translation pipeline approach, translating from the source language to English, summarizing and translating back. Evaluating with five different metrics, we find that there is variation across LLMs in their performance across similar parameter sizes, that our multilingual fine-tuned mT5 baseline outperforms most other approaches including zero-shot LLM performance for most metrics, and that LLM as judge may be less reliable on less-resourced languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comparing Approaches to Automatic Summarization in Less-Resourced Languages
Palen-Michel, Chester
Lignos, Constantine
Computation and Language
Artificial Intelligence
Automatic text summarization has achieved high performance in high-resourced languages like English, but comparatively less attention has been given to summarization in less-resourced languages. This work compares a variety of different approaches to summarization from zero-shot prompting of LLMs large and small to fine-tuning smaller models like mT5 with and without three data augmentation approaches and multilingual transfer. We also explore an LLM translation pipeline approach, translating from the source language to English, summarizing and translating back. Evaluating with five different metrics, we find that there is variation across LLMs in their performance across similar parameter sizes, that our multilingual fine-tuned mT5 baseline outperforms most other approaches including zero-shot LLM performance for most metrics, and that LLM as judge may be less reliable on less-resourced languages.
title Comparing Approaches to Automatic Summarization in Less-Resourced Languages
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.24410