Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goel, Anmol, Emde, Cornelius, Yun, Sangdoo, Oh, Seong Joon, Gubri, Martin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914490677723136
author Goel, Anmol
Emde, Cornelius
Yun, Sangdoo
Oh, Seong Joon
Gubri, Martin
author_facet Goel, Anmol
Emde, Cornelius
Yun, Sangdoo
Oh, Seong Joon
Gubri, Martin
contents We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code printing internal variables, among others. Fine-tuned models lose their ability to reason about contextual privacy norms, share information inappropriately with tools, and violate memory boundaries across contexts. Privacy collapse is a ``silent failure'' because models maintain high performance on standard safety and utility benchmarks whilst exhibiting severe privacy vulnerabilities. Our experiments show evidence of privacy collapse across six models (closed and open weight), five fine-tuning datasets (real-world and controlled data), and two task categories (agentic and memory-based). Our mechanistic analysis reveals that privacy representations are uniquely fragile to fine-tuning, compared to task-relevant features which are preserved. Our results reveal a critical gap in current safety evaluations, in particular for the deployment of specialised agents.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15220
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
Goel, Anmol
Emde, Cornelius
Yun, Sangdoo
Oh, Seong Joon
Gubri, Martin
Computation and Language
We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle patterns in training data can degrade contextual privacy, including optimisation for helpfulness, exposure to user information, emotional and subjective dialogue, and debugging code printing internal variables, among others. Fine-tuned models lose their ability to reason about contextual privacy norms, share information inappropriately with tools, and violate memory boundaries across contexts. Privacy collapse is a ``silent failure'' because models maintain high performance on standard safety and utility benchmarks whilst exhibiting severe privacy vulnerabilities. Our experiments show evidence of privacy collapse across six models (closed and open weight), five fine-tuning datasets (real-world and controlled data), and two task categories (agentic and memory-based). Our mechanistic analysis reveals that privacy representations are uniquely fragile to fine-tuning, compared to task-relevant features which are preserved. Our results reveal a critical gap in current safety evaluations, in particular for the deployment of specialised agents.
title Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models
topic Computation and Language
url https://arxiv.org/abs/2601.15220