AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915686047023104 |
|---|---|
| author | Gaikwad, Madhava |
| author_facet | Gaikwad, Madhava |
| contents | Large language models are exposed to risks of extraction, distillation, and unauthorized fine-tuning. Existing defenses use watermarking or monitoring, but these act after leakage. We design AlignDP, a hybrid privacy lock that blocks knowledge transfer at the data interface. The key idea is to separate rare and non-rare fields. Rare fields are shielded by PAC indistinguishability, giving effective zero-epsilon local DP. Non-rare fields are privatized with RAPPOR, giving unbiased frequency estimates under local DP. A global aggregator enforces composition and budget. This two-tier design hides rare events and adds controlled noise to frequent events. We prove limits of PAC extension to global aggregation, give bounds for RAPPOR estimates, and analyze utility trade-off. A toy simulation confirms feasibility: rare categories remain hidden, frequent categories are recovered with small error. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_17251 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs Gaikwad, Madhava Cryptography and Security Artificial Intelligence Machine Learning 68T05 H.2.7; I.2.7; K.4.1 Large language models are exposed to risks of extraction, distillation, and unauthorized fine-tuning. Existing defenses use watermarking or monitoring, but these act after leakage. We design AlignDP, a hybrid privacy lock that blocks knowledge transfer at the data interface. The key idea is to separate rare and non-rare fields. Rare fields are shielded by PAC indistinguishability, giving effective zero-epsilon local DP. Non-rare fields are privatized with RAPPOR, giving unbiased frequency estimates under local DP. A global aggregator enforces composition and budget. This two-tier design hides rare events and adds controlled noise to frequent events. We prove limits of PAC extension to global aggregation, give bounds for RAPPOR estimates, and analyze utility trade-off. A toy simulation confirms feasibility: rare categories remain hidden, frequent categories are recovered with small error. |
| title | AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs |
| topic | Cryptography and Security Artificial Intelligence Machine Learning 68T05 H.2.7; I.2.7; K.4.1 |
| url | https://arxiv.org/abs/2512.17251 |