Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yodthapa, Nathikan, Intharah, Thanapong, Bulathwela, Sahan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914491475689472
author Yodthapa, Nathikan
Intharah, Thanapong
Bulathwela, Sahan
author_facet Yodthapa, Nathikan
Intharah, Thanapong
Bulathwela, Sahan
contents Topic-controlled summarisation enables users to generate summaries focused on specific aspects of source documents. This paper investigates a data augmentation strategy for training small language models (sLMs) to perform topic-controlled summarisation. We propose a pairwise data augmentation method that combines contexts from different documents to create contrastive training examples, enabling models to learn the relationship between topics and summaries more effectively. Using the SciTLDR dataset enriched with Wikipedia-derived topics, we systematically evaluate how augmentation scale affects model performance. Results show consistent improvements in win rate and semantic alignment as the augmentation scale increases, while the amount of real training data remains fixed. Consequently, a T5-base model trained with our augmentation approach achieves competitive performance relative to larger models, despite using significantly fewer parameters and substantially fewer real training examples.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18087
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation
Yodthapa, Nathikan
Intharah, Thanapong
Bulathwela, Sahan
Computation and Language
Artificial Intelligence
Computers and Society
H.3.3; J.1; I.2.0
Topic-controlled summarisation enables users to generate summaries focused on specific aspects of source documents. This paper investigates a data augmentation strategy for training small language models (sLMs) to perform topic-controlled summarisation. We propose a pairwise data augmentation method that combines contexts from different documents to create contrastive training examples, enabling models to learn the relationship between topics and summaries more effectively. Using the SciTLDR dataset enriched with Wikipedia-derived topics, we systematically evaluate how augmentation scale affects model performance. Results show consistent improvements in win rate and semantic alignment as the augmentation scale increases, while the amount of real training data remains fixed. Consequently, a T5-base model trained with our augmentation approach achieves competitive performance relative to larger models, despite using significantly fewer parameters and substantially fewer real training examples.
title Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation
topic Computation and Language
Artificial Intelligence
Computers and Society
H.3.3; J.1; I.2.0
url https://arxiv.org/abs/2604.18087