Critical Learning Periods: Leveraging Early Training Dynamics for Efficient Data Pruning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chimoto, Everlyn Asiko, Gala, Jay, Ahia, Orevaoghene, Kreutzer, Julia, Bassett, Bruce A., Hooker, Sara
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911928392089600
author Chimoto, Everlyn Asiko
Gala, Jay
Ahia, Orevaoghene
Kreutzer, Julia
Bassett, Bruce A.
Hooker, Sara
author_facet Chimoto, Everlyn Asiko
Gala, Jay
Ahia, Orevaoghene
Kreutzer, Julia
Bassett, Bruce A.
Hooker, Sara
contents Neural Machine Translation models are extremely data and compute-hungry. However, not all data points contribute equally to model training and generalization. Data pruning to remove the low-value data points has the benefit of drastically reducing the compute budget without significant drop in model performance. In this paper, we propose a new data pruning technique: Checkpoints Across Time (CAT), that leverages early model training dynamics to identify the most relevant data points for model performance. We benchmark CAT against several data pruning techniques including COMET-QE, LASER and LaBSE. We find that CAT outperforms the benchmarks on Indo-European languages on multiple test sets. When applied to English-German, English-French and English-Swahili translation tasks, CAT achieves comparable performance to using the full dataset, while pruning up to 50% of training data. We inspect the data points that CAT selects and find that it tends to favour longer sentences and sentences with unique or rare words.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19462
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Critical Learning Periods: Leveraging Early Training Dynamics for Efficient Data Pruning
Chimoto, Everlyn Asiko
Gala, Jay
Ahia, Orevaoghene
Kreutzer, Julia
Bassett, Bruce A.
Hooker, Sara
Computation and Language
Neural Machine Translation models are extremely data and compute-hungry. However, not all data points contribute equally to model training and generalization. Data pruning to remove the low-value data points has the benefit of drastically reducing the compute budget without significant drop in model performance. In this paper, we propose a new data pruning technique: Checkpoints Across Time (CAT), that leverages early model training dynamics to identify the most relevant data points for model performance. We benchmark CAT against several data pruning techniques including COMET-QE, LASER and LaBSE. We find that CAT outperforms the benchmarks on Indo-European languages on multiple test sets. When applied to English-German, English-French and English-Swahili translation tasks, CAT achieves comparable performance to using the full dataset, while pruning up to 50% of training data. We inspect the data points that CAT selects and find that it tends to favour longer sentences and sentences with unique or rare words.
title Critical Learning Periods: Leveraging Early Training Dynamics for Efficient Data Pruning
topic Computation and Language
url https://arxiv.org/abs/2405.19462