Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tianjian, Xu, Haoran, Koehn, Philipp, Khashabi, Daniel, Murray, Kenton
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909139925467136
author Li, Tianjian
Xu, Haoran
Koehn, Philipp
Khashabi, Daniel
Murray, Kenton
author_facet Li, Tianjian
Xu, Haoran
Koehn, Philipp
Khashabi, Daniel
Murray, Kenton
contents Text generation models are notoriously vulnerable to errors in the training data. With the wide-spread availability of massive amounts of web-crawled data becoming more commonplace, how can we enhance the robustness of models trained on a massive amount of noisy web-crawled text? In our work, we propose Error Norm Truncation (ENT), a robust enhancement method to the standard training objective that truncates noisy data. Compared to methods that only uses the negative log-likelihood loss to estimate data quality, our method provides a more accurate estimation by considering the distribution of non-target tokens, which is often overlooked by previous work. Through comprehensive experiments across language modeling, machine translation, and text summarization, we show that equipping text generation models with ENT improves generation quality over standard training and previous soft and hard truncation methods. Furthermore, we show that our method improves the robustness of models against two of the most detrimental types of noise in machine translation, resulting in an increase of more than 2 BLEU points over the MLE baseline when up to 50% of noise is added to the data.
format Preprint
id arxiv_https___arxiv_org_abs_2310_00840
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
Li, Tianjian
Xu, Haoran
Koehn, Philipp
Khashabi, Daniel
Murray, Kenton
Computation and Language
Text generation models are notoriously vulnerable to errors in the training data. With the wide-spread availability of massive amounts of web-crawled data becoming more commonplace, how can we enhance the robustness of models trained on a massive amount of noisy web-crawled text? In our work, we propose Error Norm Truncation (ENT), a robust enhancement method to the standard training objective that truncates noisy data. Compared to methods that only uses the negative log-likelihood loss to estimate data quality, our method provides a more accurate estimation by considering the distribution of non-target tokens, which is often overlooked by previous work. Through comprehensive experiments across language modeling, machine translation, and text summarization, we show that equipping text generation models with ENT improves generation quality over standard training and previous soft and hard truncation methods. Furthermore, we show that our method improves the robustness of models against two of the most detrimental types of noise in machine translation, resulting in an increase of more than 2 BLEU points over the MLE baseline when up to 50% of noise is added to the data.
title Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation Models
topic Computation and Language
url https://arxiv.org/abs/2310.00840