ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Tong, Brahman, Faeze, Liu, Jiacheng, Mireshghallah, Niloofar, Shi, Weijia, Koh, Pang Wei, Zettlemoyer, Luke, Hajishirzi, Hannaneh
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911063034822656
author Chen, Tong
Brahman, Faeze
Liu, Jiacheng
Mireshghallah, Niloofar
Shi, Weijia
Koh, Pang Wei
Zettlemoyer, Luke
Hajishirzi, Hannaneh
author_facet Chen, Tong
Brahman, Faeze
Liu, Jiacheng
Mireshghallah, Niloofar
Shi, Weijia
Koh, Pang Wei
Zettlemoyer, Luke
Hajishirzi, Hannaneh
contents Language models (LMs) can memorize and reproduce segments from their pretraining data verbatim even in non-adversarial settings, raising concerns about copyright, plagiarism, privacy, and creativity. We introduce Paraphrase Preference Optimization (ParaPO), a post-training method that fine-tunes LMs to reduce unintentional regurgitation while preserving their overall utility. ParaPO trains LMs to prefer paraphrased versions of memorized segments over the original verbatim content from the pretraining data. To maintain the ability to recall famous quotations when appropriate, we develop a variant of ParaPO that uses system prompts to control regurgitation behavior. In our evaluation on Llama3.1-8B, ParaPO consistently reduces regurgitation across all tested datasets (e.g., reducing the regurgitation metric from 17.3 to 12.9 in creative writing), whereas unlearning methods used in prior work to mitigate regurgitation are less effective outside their targeted unlearned domain (from 17.3 to 16.9). When applied to the instruction-tuned Tulu3-8B model, ParaPO with system prompting successfully preserves famous quotation recall while reducing unintentional regurgitation (from 8.7 to 6.3 in creative writing) when prompted not to regurgitate. In contrast, without ParaPO tuning, prompting the model not to regurgitate produces only a marginal reduction (8.7 to 8.4).
format Preprint
id arxiv_https___arxiv_org_abs_2504_14452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
Chen, Tong
Brahman, Faeze
Liu, Jiacheng
Mireshghallah, Niloofar
Shi, Weijia
Koh, Pang Wei
Zettlemoyer, Luke
Hajishirzi, Hannaneh
Computation and Language
Artificial Intelligence
Machine Learning
Language models (LMs) can memorize and reproduce segments from their pretraining data verbatim even in non-adversarial settings, raising concerns about copyright, plagiarism, privacy, and creativity. We introduce Paraphrase Preference Optimization (ParaPO), a post-training method that fine-tunes LMs to reduce unintentional regurgitation while preserving their overall utility. ParaPO trains LMs to prefer paraphrased versions of memorized segments over the original verbatim content from the pretraining data. To maintain the ability to recall famous quotations when appropriate, we develop a variant of ParaPO that uses system prompts to control regurgitation behavior. In our evaluation on Llama3.1-8B, ParaPO consistently reduces regurgitation across all tested datasets (e.g., reducing the regurgitation metric from 17.3 to 12.9 in creative writing), whereas unlearning methods used in prior work to mitigate regurgitation are less effective outside their targeted unlearned domain (from 17.3 to 16.9). When applied to the instruction-tuned Tulu3-8B model, ParaPO with system prompting successfully preserves famous quotation recall while reducing unintentional regurgitation (from 8.7 to 6.3 in creative writing) when prompted not to regurgitate. In contrast, without ParaPO tuning, prompting the model not to regurgitate produces only a marginal reduction (8.7 to 8.4).
title ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.14452