PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hughes, Anthony, Duddu, Vasisht, Asokan, N., Aletras, Nikolaos, Ma, Ning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917296372449280
author Hughes, Anthony
Duddu, Vasisht
Asokan, N.
Aletras, Nikolaos
Ma, Ning
author_facet Hughes, Anthony
Duddu, Vasisht
Asokan, N.
Aletras, Nikolaos
Ma, Ning
contents Language models (LMs) may memorize personally identifiable information (PII) from training data, enabling adversaries to extract it during inference. Existing defense mechanisms such as differential privacy (DP) reduce this leakage, but incur large drops in utility. Based on a comprehensive study using circuit discovery to identify the computational circuits responsible PII leakage in LMs, we hypothesize that specific PII leakage circuits in LMs should be responsible for this behavior. Therefore, we propose PATCH (Privacy-Aware Targeted Circuit PatcHing), a novel approach that first identifies and subsequently directly edits PII circuits to reduce leakage. PATCH achieves better privacy-utility trade-off than existing defenses, e.g., reducing recall of PII leakage from LMs by up to 65%. Finally, PATCH can be combined with DP to reduce recall of residual leakage of an LM to as low as 0.01%. Our analysis shows that PII leakage circuits persist even after the application of existing defense mechanisms. In contrast, PATCH can effectively mitigate their impact.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
Hughes, Anthony
Duddu, Vasisht
Asokan, N.
Aletras, Nikolaos
Ma, Ning
Cryptography and Security
Computation and Language
Language models (LMs) may memorize personally identifiable information (PII) from training data, enabling adversaries to extract it during inference. Existing defense mechanisms such as differential privacy (DP) reduce this leakage, but incur large drops in utility. Based on a comprehensive study using circuit discovery to identify the computational circuits responsible PII leakage in LMs, we hypothesize that specific PII leakage circuits in LMs should be responsible for this behavior. Therefore, we propose PATCH (Privacy-Aware Targeted Circuit PatcHing), a novel approach that first identifies and subsequently directly edits PII circuits to reduce leakage. PATCH achieves better privacy-utility trade-off than existing defenses, e.g., reducing recall of PII leakage from LMs by up to 65%. Finally, PATCH can be combined with DP to reduce recall of residual leakage of an LM to as low as 0.01%. Our analysis shows that PII leakage circuits persist even after the application of existing defense mechanisms. In contrast, PATCH can effectively mitigate their impact.
title PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
topic Cryptography and Security
Computation and Language
url https://arxiv.org/abs/2510.07452