SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Nan, Xiong, Weichen, Liu, Hanwen, Liao, Yi, Ding, Lei, Zhang, Kai, Tang, Guohua, Han, Xiao, Yang, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909248222396416
author He, Nan
Xiong, Weichen
Liu, Hanwen
Liao, Yi
Ding, Lei
Zhang, Kai
Tang, Guohua
Han, Xiao
Yang, Wei
author_facet He, Nan
Xiong, Weichen
Liu, Hanwen
Liao, Yi
Ding, Lei
Zhang, Kai
Tang, Guohua
Han, Xiao
Yang, Wei
contents The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. To address this, we propose a soft deduplication method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. Central to our approach is the concept of "data commonness", a metric we introduce to quantify the degree of duplication by measuring the occurrence probabilities of samples using an n-gram model. Empirical analysis shows that this method significantly improves training efficiency, achieving comparable perplexity scores with at least a 26% reduction in required training steps. Additionally, it enhances average few-shot downstream accuracy by 1.77% when trained for an equivalent duration. Importantly, this approach consistently improves performance, even on rigorously deduplicated datasets, indicating its potential to complement existing methods and become a standard pre-training process for LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06654
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
He, Nan
Xiong, Weichen
Liu, Hanwen
Liao, Yi
Ding, Lei
Zhang, Kai
Tang, Guohua
Han, Xiao
Yang, Wei
Computation and Language
Artificial Intelligence
The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. To address this, we propose a soft deduplication method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. Central to our approach is the concept of "data commonness", a metric we introduce to quantify the degree of duplication by measuring the occurrence probabilities of samples using an n-gram model. Empirical analysis shows that this method significantly improves training efficiency, achieving comparable perplexity scores with at least a 26% reduction in required training steps. Additionally, it enhances average few-shot downstream accuracy by 1.77% when trained for an equivalent duration. Importantly, this approach consistently improves performance, even on rigorously deduplicated datasets, indicating its potential to complement existing methods and become a standard pre-training process for LLMs.
title SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.06654