Towards Confidential and Efficient LLM Inference with Dual Privacy Protection
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916946244534272 |
|---|---|
| author | Yu, Honglan Wang, Yibin Dai, Feifei Liu, Dong Fan, Haihui Gu, Xiaoyan |
| author_facet | Yu, Honglan Wang, Yibin Dai, Feifei Liu, Dong Fan, Haihui Gu, Xiaoyan |
| contents | CPU-based trusted execution environments (TEEs) and differential privacy (DP) have gained wide applications for private inference. Due to high inference latency in TEEs, researchers use partition-based approaches that offload linear model components to GPUs. However, dense nonlinear layers of large language models (LLMs) result in significant communication overhead between TEEs and GPUs. DP-based approaches apply random noise to protect data privacy, but this compromises LLM performance and semantic understanding. To overcome the above drawbacks, this paper proposes CMIF, a Confidential and efficient Model Inference Framework. CMIF confidentially deploys the embedding layer in the client-side TEE and subsequent layers on GPU servers. Meanwhile, it optimizes the Report-Noisy-Max mechanism to protect sensitive inputs with a slight decrease in model performance. Extensive experiments on Llama-series models demonstrate that CMIF reduces additional inference overhead in TEEs while preserving user data privacy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_09091 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Confidential and Efficient LLM Inference with Dual Privacy Protection Yu, Honglan Wang, Yibin Dai, Feifei Liu, Dong Fan, Haihui Gu, Xiaoyan Cryptography and Security Artificial Intelligence CPU-based trusted execution environments (TEEs) and differential privacy (DP) have gained wide applications for private inference. Due to high inference latency in TEEs, researchers use partition-based approaches that offload linear model components to GPUs. However, dense nonlinear layers of large language models (LLMs) result in significant communication overhead between TEEs and GPUs. DP-based approaches apply random noise to protect data privacy, but this compromises LLM performance and semantic understanding. To overcome the above drawbacks, this paper proposes CMIF, a Confidential and efficient Model Inference Framework. CMIF confidentially deploys the embedding layer in the client-side TEE and subsequent layers on GPU servers. Meanwhile, it optimizes the Report-Noisy-Max mechanism to protect sensitive inputs with a slight decrease in model performance. Extensive experiments on Llama-series models demonstrate that CMIF reduces additional inference overhead in TEEs while preserving user data privacy. |
| title | Towards Confidential and Efficient LLM Inference with Dual Privacy Protection |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2509.09091 |