If It's Nice, Do It Twice: We Should Try Iterative Corpus Curation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Young, Robin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911416831705088
author Young, Robin
author_facet Young, Robin
contents Recent work demonstrates that filtering harmful content from pretraining data improves model safety without degrading capabilities. We propose a natural extension: do it again. A model trained on filtered data can filter the corpus further; training on this cleaner corpus produces an even cleaner model. We provide theoretical analysis showing this process converges to a self-consistent corpus where the model trained on it approves of its own training data. Even under the weak assumption of constant filter quality, iteration yields decay in harmful content. We argue this framework offers a novel form of scalable oversight. While model internals are opaque, the resulting corpus is human-auditable. Even a single iteration produces a large-scale preference annotations over documents, potentially valuable for interpretability research. We derive bounds on capability-safety tradeoffs and outline open questions. We call on researchers with pretraining infrastructure to empirically test this approach.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15280
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle If It's Nice, Do It Twice: We Should Try Iterative Corpus Curation
Young, Robin
Artificial Intelligence
Computers and Society
Computer Science and Game Theory
Recent work demonstrates that filtering harmful content from pretraining data improves model safety without degrading capabilities. We propose a natural extension: do it again. A model trained on filtered data can filter the corpus further; training on this cleaner corpus produces an even cleaner model. We provide theoretical analysis showing this process converges to a self-consistent corpus where the model trained on it approves of its own training data. Even under the weak assumption of constant filter quality, iteration yields decay in harmful content. We argue this framework offers a novel form of scalable oversight. While model internals are opaque, the resulting corpus is human-auditable. Even a single iteration produces a large-scale preference annotations over documents, potentially valuable for interpretability research. We derive bounds on capability-safety tradeoffs and outline open questions. We call on researchers with pretraining infrastructure to empirically test this approach.
title If It's Nice, Do It Twice: We Should Try Iterative Corpus Curation
topic Artificial Intelligence
Computers and Society
Computer Science and Game Theory
url https://arxiv.org/abs/2501.15280