SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909248222396416 |
|---|---|
| author | He, Nan Xiong, Weichen Liu, Hanwen Liao, Yi Ding, Lei Zhang, Kai Tang, Guohua Han, Xiao Yang, Wei |
| author_facet | He, Nan Xiong, Weichen Liu, Hanwen Liao, Yi Ding, Lei Zhang, Kai Tang, Guohua Han, Xiao Yang, Wei |
| contents | The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. To address this, we propose a soft deduplication method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. Central to our approach is the concept of "data commonness", a metric we introduce to quantify the degree of duplication by measuring the occurrence probabilities of samples using an n-gram model. Empirical analysis shows that this method significantly improves training efficiency, achieving comparable perplexity scores with at least a 26% reduction in required training steps. Additionally, it enhances average few-shot downstream accuracy by 1.77% when trained for an equivalent duration. Importantly, this approach consistently improves performance, even on rigorously deduplicated datasets, indicating its potential to complement existing methods and become a standard pre-training process for LLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_06654 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training He, Nan Xiong, Weichen Liu, Hanwen Liao, Yi Ding, Lei Zhang, Kai Tang, Guohua Han, Xiao Yang, Wei Computation and Language Artificial Intelligence The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. To address this, we propose a soft deduplication method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. Central to our approach is the concept of "data commonness", a metric we introduce to quantify the degree of duplication by measuring the occurrence probabilities of samples using an n-gram model. Empirical analysis shows that this method significantly improves training efficiency, achieving comparable perplexity scores with at least a 26% reduction in required training steps. Additionally, it enhances average few-shot downstream accuracy by 1.77% when trained for an equivalent duration. Importantly, this approach consistently improves performance, even on rigorously deduplicated datasets, indicating its potential to complement existing methods and become a standard pre-training process for LLMs. |
| title | SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2407.06654 |