Noise-Aware Training of Layout-Aware Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sarkhel, Ritesh, Ren, Xiaoqi, Costa, Lauro Beltrao, Su, Guolong, Perot, Vincent, Xie, Yanan, Koukoumidis, Emmanouil, Nandi, Arnab
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910392807063552
author Sarkhel, Ritesh
Ren, Xiaoqi
Costa, Lauro Beltrao
Su, Guolong
Perot, Vincent
Xie, Yanan
Koukoumidis, Emmanouil
Nandi, Arnab
author_facet Sarkhel, Ritesh
Ren, Xiaoqi
Costa, Lauro Beltrao
Su, Guolong
Perot, Vincent
Xie, Yanan
Koukoumidis, Emmanouil
Nandi, Arnab
contents A visually rich document (VRD) utilizes visual features along with linguistic cues to disseminate information. Training a custom extractor that identifies named entities from a document requires a large number of instances of the target document type annotated at textual and visual modalities. This is an expensive bottleneck in enterprise scenarios, where we want to train custom extractors for thousands of different document types in a scalable way. Pre-training an extractor model on unlabeled instances of the target document type, followed by a fine-tuning step on human-labeled instances does not work in these scenarios, as it surpasses the maximum allowable training time allocated for the extractor. We address this scenario by proposing a Noise-Aware Training method or NAT in this paper. Instead of acquiring expensive human-labeled documents, NAT utilizes weakly labeled documents to train an extractor in a scalable way. To avoid degradation in the model's quality due to noisy, weakly labeled samples, NAT estimates the confidence of each training sample and incorporates it as uncertainty measure during training. We train multiple state-of-the-art extractor models using NAT. Experiments on a number of publicly available and in-house datasets show that NAT-trained models are not only robust in performance -- it outperforms a transfer-learning baseline by up to 6% in terms of macro-F1 score, but it is also more label-efficient -- it reduces the amount of human-effort required to obtain comparable performance by up to 73%.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00488
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Noise-Aware Training of Layout-Aware Language Models
Sarkhel, Ritesh
Ren, Xiaoqi
Costa, Lauro Beltrao
Su, Guolong
Perot, Vincent
Xie, Yanan
Koukoumidis, Emmanouil
Nandi, Arnab
Computation and Language
Artificial Intelligence
Machine Learning
A visually rich document (VRD) utilizes visual features along with linguistic cues to disseminate information. Training a custom extractor that identifies named entities from a document requires a large number of instances of the target document type annotated at textual and visual modalities. This is an expensive bottleneck in enterprise scenarios, where we want to train custom extractors for thousands of different document types in a scalable way. Pre-training an extractor model on unlabeled instances of the target document type, followed by a fine-tuning step on human-labeled instances does not work in these scenarios, as it surpasses the maximum allowable training time allocated for the extractor. We address this scenario by proposing a Noise-Aware Training method or NAT in this paper. Instead of acquiring expensive human-labeled documents, NAT utilizes weakly labeled documents to train an extractor in a scalable way. To avoid degradation in the model's quality due to noisy, weakly labeled samples, NAT estimates the confidence of each training sample and incorporates it as uncertainty measure during training. We train multiple state-of-the-art extractor models using NAT. Experiments on a number of publicly available and in-house datasets show that NAT-trained models are not only robust in performance -- it outperforms a transfer-learning baseline by up to 6% in terms of macro-F1 score, but it is also more label-efficient -- it reduces the amount of human-effort required to obtain comparable performance by up to 73%.
title Noise-Aware Training of Layout-Aware Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.00488