Towards Confidential and Efficient LLM Inference with Dual Privacy Protection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Honglan, Wang, Yibin, Dai, Feifei, Liu, Dong, Fan, Haihui, Gu, Xiaoyan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916946244534272
author Yu, Honglan
Wang, Yibin
Dai, Feifei
Liu, Dong
Fan, Haihui
Gu, Xiaoyan
author_facet Yu, Honglan
Wang, Yibin
Dai, Feifei
Liu, Dong
Fan, Haihui
Gu, Xiaoyan
contents CPU-based trusted execution environments (TEEs) and differential privacy (DP) have gained wide applications for private inference. Due to high inference latency in TEEs, researchers use partition-based approaches that offload linear model components to GPUs. However, dense nonlinear layers of large language models (LLMs) result in significant communication overhead between TEEs and GPUs. DP-based approaches apply random noise to protect data privacy, but this compromises LLM performance and semantic understanding. To overcome the above drawbacks, this paper proposes CMIF, a Confidential and efficient Model Inference Framework. CMIF confidentially deploys the embedding layer in the client-side TEE and subsequent layers on GPU servers. Meanwhile, it optimizes the Report-Noisy-Max mechanism to protect sensitive inputs with a slight decrease in model performance. Extensive experiments on Llama-series models demonstrate that CMIF reduces additional inference overhead in TEEs while preserving user data privacy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09091
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Confidential and Efficient LLM Inference with Dual Privacy Protection
Yu, Honglan
Wang, Yibin
Dai, Feifei
Liu, Dong
Fan, Haihui
Gu, Xiaoyan
Cryptography and Security
Artificial Intelligence
CPU-based trusted execution environments (TEEs) and differential privacy (DP) have gained wide applications for private inference. Due to high inference latency in TEEs, researchers use partition-based approaches that offload linear model components to GPUs. However, dense nonlinear layers of large language models (LLMs) result in significant communication overhead between TEEs and GPUs. DP-based approaches apply random noise to protect data privacy, but this compromises LLM performance and semantic understanding. To overcome the above drawbacks, this paper proposes CMIF, a Confidential and efficient Model Inference Framework. CMIF confidentially deploys the embedding layer in the client-side TEE and subsequent layers on GPU servers. Meanwhile, it optimizes the Report-Noisy-Max mechanism to protect sensitive inputs with a slight decrease in model performance. Extensive experiments on Llama-series models demonstrate that CMIF reduces additional inference overhead in TEEs while preserving user data privacy.
title Towards Confidential and Efficient LLM Inference with Dual Privacy Protection
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2509.09091