Self-Compositional Data Augmentation for Scientific Keyphrase Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866929578692313088 |
|---|---|
| author | Houbre, Mael Boudin, Florian Daille, Beatrice Aizawa, Akiko |
| author_facet | Houbre, Mael Boudin, Florian Daille, Beatrice Aizawa, Akiko |
| contents | State-of-the-art models for keyphrase generation require large amounts of training data to achieve good performance. However, obtaining keyphrase-labeled documents can be challenging and costly. To address this issue, we present a self-compositional data augmentation method. More specifically, we measure the relatedness of training documents based on their shared keyphrases, and combine similar documents to generate synthetic samples. The advantage of our method lies in its ability to create additional training samples that keep domain coherence, without relying on external data or resources. Our results on multiple datasets spanning three different domains, demonstrate that our method consistently improves keyphrase generation. A qualitative analysis of the generated keyphrases for the Computer Science domain confirms this improvement towards their representativity property. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_03039 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Self-Compositional Data Augmentation for Scientific Keyphrase Generation Houbre, Mael Boudin, Florian Daille, Beatrice Aizawa, Akiko Computation and Language Information Retrieval State-of-the-art models for keyphrase generation require large amounts of training data to achieve good performance. However, obtaining keyphrase-labeled documents can be challenging and costly. To address this issue, we present a self-compositional data augmentation method. More specifically, we measure the relatedness of training documents based on their shared keyphrases, and combine similar documents to generate synthetic samples. The advantage of our method lies in its ability to create additional training samples that keep domain coherence, without relying on external data or resources. Our results on multiple datasets spanning three different domains, demonstrate that our method consistently improves keyphrase generation. A qualitative analysis of the generated keyphrases for the Computer Science domain confirms this improvement towards their representativity property. |
| title | Self-Compositional Data Augmentation for Scientific Keyphrase Generation |
| topic | Computation and Language Information Retrieval |
| url | https://arxiv.org/abs/2411.03039 |