A Bitter Lesson for Data Filtering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mohri, Christopher, Duchi, John, Hashimoto, Tatsunori
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917510228475904
author Mohri, Christopher
Duchi, John
Hashimoto, Tatsunori
author_facet Mohri, Christopher
Duchi, John
Hashimoto, Tatsunori
contents We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally ``poor'' data.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19407
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Bitter Lesson for Data Filtering
Mohri, Christopher
Duchi, John
Hashimoto, Tatsunori
Machine Learning
Artificial Intelligence
We investigate data filtering for large model pretraining via new scaling studies that target the high compute, data-scarce regime. In spite of an apparently common belief that filtering data to include only high-quality information is essential, our experiments suggest that with enough compute, the best data filter is no data filter. We find that sufficiently trained large parameter models not only tolerate low-quality and distractor data, but in fact benefit from nominally ``poor'' data.
title A Bitter Lesson for Data Filtering
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.19407