AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Gaikwad, Madhava
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915686047023104
author Gaikwad, Madhava
author_facet Gaikwad, Madhava
contents Large language models are exposed to risks of extraction, distillation, and unauthorized fine-tuning. Existing defenses use watermarking or monitoring, but these act after leakage. We design AlignDP, a hybrid privacy lock that blocks knowledge transfer at the data interface. The key idea is to separate rare and non-rare fields. Rare fields are shielded by PAC indistinguishability, giving effective zero-epsilon local DP. Non-rare fields are privatized with RAPPOR, giving unbiased frequency estimates under local DP. A global aggregator enforces composition and budget. This two-tier design hides rare events and adds controlled noise to frequent events. We prove limits of PAC extension to global aggregation, give bounds for RAPPOR estimates, and analyze utility trade-off. A toy simulation confirms feasibility: rare categories remain hidden, frequent categories are recovered with small error.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17251
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
Gaikwad, Madhava
Cryptography and Security
Artificial Intelligence
Machine Learning
68T05
H.2.7; I.2.7; K.4.1
Large language models are exposed to risks of extraction, distillation, and unauthorized fine-tuning. Existing defenses use watermarking or monitoring, but these act after leakage. We design AlignDP, a hybrid privacy lock that blocks knowledge transfer at the data interface. The key idea is to separate rare and non-rare fields. Rare fields are shielded by PAC indistinguishability, giving effective zero-epsilon local DP. Non-rare fields are privatized with RAPPOR, giving unbiased frequency estimates under local DP. A global aggregator enforces composition and budget. This two-tier design hides rare events and adds controlled noise to frequent events. We prove limits of PAC extension to global aggregation, give bounds for RAPPOR estimates, and analyze utility trade-off. A toy simulation confirms feasibility: rare categories remain hidden, frequent categories are recovered with small error.
title AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
topic Cryptography and Security
Artificial Intelligence
Machine Learning
68T05
H.2.7; I.2.7; K.4.1
url https://arxiv.org/abs/2512.17251