Self-Compositional Data Augmentation for Scientific Keyphrase Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Houbre, Mael, Boudin, Florian, Daille, Beatrice, Aizawa, Akiko
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929578692313088
author Houbre, Mael
Boudin, Florian
Daille, Beatrice
Aizawa, Akiko
author_facet Houbre, Mael
Boudin, Florian
Daille, Beatrice
Aizawa, Akiko
contents State-of-the-art models for keyphrase generation require large amounts of training data to achieve good performance. However, obtaining keyphrase-labeled documents can be challenging and costly. To address this issue, we present a self-compositional data augmentation method. More specifically, we measure the relatedness of training documents based on their shared keyphrases, and combine similar documents to generate synthetic samples. The advantage of our method lies in its ability to create additional training samples that keep domain coherence, without relying on external data or resources. Our results on multiple datasets spanning three different domains, demonstrate that our method consistently improves keyphrase generation. A qualitative analysis of the generated keyphrases for the Computer Science domain confirms this improvement towards their representativity property.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Compositional Data Augmentation for Scientific Keyphrase Generation
Houbre, Mael
Boudin, Florian
Daille, Beatrice
Aizawa, Akiko
Computation and Language
Information Retrieval
State-of-the-art models for keyphrase generation require large amounts of training data to achieve good performance. However, obtaining keyphrase-labeled documents can be challenging and costly. To address this issue, we present a self-compositional data augmentation method. More specifically, we measure the relatedness of training documents based on their shared keyphrases, and combine similar documents to generate synthetic samples. The advantage of our method lies in its ability to create additional training samples that keep domain coherence, without relying on external data or resources. Our results on multiple datasets spanning three different domains, demonstrate that our method consistently improves keyphrase generation. A qualitative analysis of the generated keyphrases for the Computer Science domain confirms this improvement towards their representativity property.
title Self-Compositional Data Augmentation for Scientific Keyphrase Generation
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2411.03039