Data Augmentation Techniques for Process Extraction from Scientific Publications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Susanti, Yuni
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915242225696768
author Susanti, Yuni
author_facet Susanti, Yuni
contents We present data augmentation techniques for process extraction tasks in scientific publications. We cast the process extraction task as a sequence labeling task where we identify all the entities in a sentence and label them according to their process-specific roles. The proposed method attempts to create meaningful augmented sentences by utilizing (1) process-specific information from the original sentence, (2) role label similarity, and (3) sentence similarity. We demonstrate that the proposed methods substantially improve the performance of the process extraction model trained on chemistry domain datasets, up to 12.3 points improvement in performance accuracy (F-score). The proposed methods could potentially reduce overfitting as well, especially when training on small datasets or in a low-resource setting such as in chemistry and other scientific domains.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Augmentation Techniques for Process Extraction from Scientific Publications
Susanti, Yuni
Computation and Language
Information Retrieval
We present data augmentation techniques for process extraction tasks in scientific publications. We cast the process extraction task as a sequence labeling task where we identify all the entities in a sentence and label them according to their process-specific roles. The proposed method attempts to create meaningful augmented sentences by utilizing (1) process-specific information from the original sentence, (2) role label similarity, and (3) sentence similarity. We demonstrate that the proposed methods substantially improve the performance of the process extraction model trained on chemistry domain datasets, up to 12.3 points improvement in performance accuracy (F-score). The proposed methods could potentially reduce overfitting as well, especially when training on small datasets or in a low-resource setting such as in chemistry and other scientific domains.
title Data Augmentation Techniques for Process Extraction from Scientific Publications
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2405.14594