Differentially Private KV-Cache Transmission for Split LLM Inference

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Yu, Renjie
Format: Recurso digital
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901835817680896
author Yu, Renjie
author_facet Yu, Renjie
contents <p>KV-cache inversion attacks can reconstruct sensitive user inputs from intercepted Key-Value vectors in split LLM inference. Existing defenses rely on heuristic obfuscation without formal privacy guarantees. We propose HeadDP, which injects calibrated Gaussian noise at the per-head key and value projection outputs to provide provable (ε, δ)-differential privacy.</p> <p>We show that naïve injection at the hidden state fails because the minimum usable privacy budget satisfies ε_min ∝ √d — a consequence of Shannon's channel capacity theorem and Fano's inequality. By injecting at the per-head projection space (d_h = 128 vs. d = 3584), we reduce ε_min by more than 7×.</p> <p>Evaluated on MMLU Medical across Qwen2.5-7B, Llama-3-8B, and Mistral-7B, HeadDP achieves an accuracy drop of ≤ 5 pp at ε = 50–75, compared to a drop of > 65 pp for hidden-state injection at ε = 200.</p> <p>Added: attack evaluation against KV-cache inversion (Section 4.3), extended benchmark on reasoning/cross-domain tasks (Section 4.4), ablation studies (Section 4.5), privacy budget discussion (Section 4.6).</p> <p>Updated: all experiments now use MMLU Medical (1,212 questions) as the unified benchmark.</p> <p>Version 4: substantially revised. Added nuclear-rank MI bound, three-architecture validation, and universal lower bound via DJW minimax framework.</p> <p> </p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19365928
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Differentially Private KV-Cache Transmission for Split LLM Inference
Yu, Renjie
differential privacy
KV-cache
split inference
LLM
<p>KV-cache inversion attacks can reconstruct sensitive user inputs from intercepted Key-Value vectors in split LLM inference. Existing defenses rely on heuristic obfuscation without formal privacy guarantees. We propose HeadDP, which injects calibrated Gaussian noise at the per-head key and value projection outputs to provide provable (ε, δ)-differential privacy.</p> <p>We show that naïve injection at the hidden state fails because the minimum usable privacy budget satisfies ε_min ∝ √d — a consequence of Shannon's channel capacity theorem and Fano's inequality. By injecting at the per-head projection space (d_h = 128 vs. d = 3584), we reduce ε_min by more than 7×.</p> <p>Evaluated on MMLU Medical across Qwen2.5-7B, Llama-3-8B, and Mistral-7B, HeadDP achieves an accuracy drop of ≤ 5 pp at ε = 50–75, compared to a drop of > 65 pp for hidden-state injection at ε = 200.</p> <p>Added: attack evaluation against KV-cache inversion (Section 4.3), extended benchmark on reasoning/cross-domain tasks (Section 4.4), ablation studies (Section 4.5), privacy budget discussion (Section 4.6).</p> <p>Updated: all experiments now use MMLU Medical (1,212 questions) as the unified benchmark.</p> <p>Version 4: substantially revised. Added nuclear-rank MI bound, three-architecture validation, and universal lower bound via DJW minimax framework.</p> <p> </p>
title Differentially Private KV-Cache Transmission for Split LLM Inference
topic differential privacy
KV-cache
split inference
LLM
url https://doi.org/10.5281/zenodo.19365928