Differentially Private and Communication Efficient Large Language Model Split Inference via Stochastic Quantization and Soft Prompt
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Yujie, Jin, Richeng, Ji, Xiaoyu, Jin, Yier, Xu, Wenyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning
von: Xia, Saisai, et al.
Veröffentlicht: (2025)
von: Xia, Saisai, et al.
Veröffentlicht: (2025)
TernaryVote: Differentially Private, Communication Efficient, and Byzantine Resilient Distributed Optimization on Heterogeneous Data
von: Jin, Richeng, et al.
Veröffentlicht: (2024)
von: Jin, Richeng, et al.
Veröffentlicht: (2024)
CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
von: Zhu, Huifeng, et al.
Veröffentlicht: (2025)
von: Zhu, Huifeng, et al.
Veröffentlicht: (2025)
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
UFO: Unlocking Ultra-Efficient Quantized Private Inference with Protocol and Algorithm Co-Optimization
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2026)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2026)
ReCIT: Reconstructing Full Private Data from Gradient in Parameter-Efficient Fine-Tuning of Large Language Models
von: Xie, Jin, et al.
Veröffentlicht: (2025)
von: Xie, Jin, et al.
Veröffentlicht: (2025)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
PRIVMARK: Private Large Language Models Watermarking with MPC
von: Fargues, Thomas, et al.
Veröffentlicht: (2025)
von: Fargues, Thomas, et al.
Veröffentlicht: (2025)
EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2024)
TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
von: Huang, Zhengxian, et al.
Veröffentlicht: (2026)
Differentially Private Parameter-Efficient Fine-tuning for Large ASR Models
von: Liu, Hongbin, et al.
Veröffentlicht: (2024)
von: Liu, Hongbin, et al.
Veröffentlicht: (2024)
Prompt Inference Attack on Distributed Large Language Model Inference Frameworks
von: Luo, Xinjian, et al.
Veröffentlicht: (2025)
von: Luo, Xinjian, et al.
Veröffentlicht: (2025)
Breaking the Communication-Privacy-Accuracy Tradeoff with $f$-Differential Privacy
von: Jin, Richeng, et al.
Veröffentlicht: (2023)
von: Jin, Richeng, et al.
Veröffentlicht: (2023)
PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts
von: Li, Qinfeng, et al.
Veröffentlicht: (2026)
von: Li, Qinfeng, et al.
Veröffentlicht: (2026)
Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models
von: Wei, Qianshan, et al.
Veröffentlicht: (2025)
von: Wei, Qianshan, et al.
Veröffentlicht: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
ConfusionPrompt: Practical Private Inference for Online Large Language Models
von: Mai, Peihua, et al.
Veröffentlicht: (2023)
von: Mai, Peihua, et al.
Veröffentlicht: (2023)
A Review and Comparison of AI Enhanced Side Channel Analysis
von: Panoff, Max, et al.
Veröffentlicht: (2024)
von: Panoff, Max, et al.
Veröffentlicht: (2024)
Enhancing Accuracy-Privacy Trade-off in Differentially Private Split Learning
von: Pham, Ngoc Duy, et al.
Veröffentlicht: (2023)
von: Pham, Ngoc Duy, et al.
Veröffentlicht: (2023)
Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
von: Shu, Dong, et al.
Veröffentlicht: (2024)
von: Shu, Dong, et al.
Veröffentlicht: (2024)
Private and Communication-Efficient Federated Learning based on Differentially Private Sketches
von: Zhang, Meifan, et al.
Veröffentlicht: (2024)
von: Zhang, Meifan, et al.
Veröffentlicht: (2024)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
von: Qu, Wenjie, et al.
Veröffentlicht: (2025)
von: Qu, Wenjie, et al.
Veröffentlicht: (2025)
DPDR: Gradient Decomposition and Reconstruction for Differentially Private Deep Learning
von: Liu, Yixuan, et al.
Veröffentlicht: (2024)
von: Liu, Yixuan, et al.
Veröffentlicht: (2024)
Invisible Finger: Practical Electromagnetic Interference Attack on Touchscreen-based Electronic Devices
von: Shan, Haoqi, et al.
Veröffentlicht: (2024)
von: Shan, Haoqi, et al.
Veröffentlicht: (2024)
Differentially Private Subspace Fine-Tuning for Large Language Models
von: Zheng, Lele, et al.
Veröffentlicht: (2026)
von: Zheng, Lele, et al.
Veröffentlicht: (2026)
PermLLM: Private Inference of Large Language Models within 3 Seconds under WAN
von: Zheng, Fei, et al.
Veröffentlicht: (2024)
von: Zheng, Fei, et al.
Veröffentlicht: (2024)
Harnessing Sparsification in Federated Learning: A Secure, Efficient, and Differentially Private Realization
von: Xu, Shuangqing, et al.
Veröffentlicht: (2025)
von: Xu, Shuangqing, et al.
Veröffentlicht: (2025)
Reconstruction of Differentially Private Text Sanitization via Large Language Models
von: Pang, Shuchao, et al.
Veröffentlicht: (2024)
von: Pang, Shuchao, et al.
Veröffentlicht: (2024)
SecPE: Secure Prompt Ensembling for Private and Robust Large Language Models
von: Zhang, Jiawen, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawen, et al.
Veröffentlicht: (2025)
ReThink: Reveal the Threat of Electromagnetic Interference on Power Inverters
von: Yang, Fengchen, et al.
Veröffentlicht: (2024)
von: Yang, Fengchen, et al.
Veröffentlicht: (2024)
Noise-Aware Differentially Private Variational Inference
von: Alrawajfeh, Talal, et al.
Veröffentlicht: (2024)
von: Alrawajfeh, Talal, et al.
Veröffentlicht: (2024)
PrivCirNet: Efficient Private Inference via Block Circulant Transformation
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity
von: Yan, Guang, et al.
Veröffentlicht: (2025)
von: Yan, Guang, et al.
Veröffentlicht: (2025)
Subsampling is not Magic: Why Large Batch Sizes Work for Differentially Private Stochastic Optimisation
von: Räisä, Ossi, et al.
Veröffentlicht: (2024)
von: Räisä, Ossi, et al.
Veröffentlicht: (2024)
Prompt Public Large Language Models to Synthesize Data for Private On-device Applications
von: Wu, Shanshan, et al.
Veröffentlicht: (2024)
von: Wu, Shanshan, et al.
Veröffentlicht: (2024)
Scaling Laws for Differentially Private Language Models
von: McKenna, Ryan, et al.
Veröffentlicht: (2025)
von: McKenna, Ryan, et al.
Veröffentlicht: (2025)
Beyond Statistical Estimation: Differentially Private Individual Computation via Shuffling
von: Wang, Shaowei, et al.
Veröffentlicht: (2024)
von: Wang, Shaowei, et al.
Veröffentlicht: (2024)
Improving Parameter-Efficient Federated Learning with Differentially Private Refactorization
von: Tran, Linh, et al.
Veröffentlicht: (2026)
von: Tran, Linh, et al.
Veröffentlicht: (2026)
DP-BREM: Differentially-Private and Byzantine-Robust Federated Learning with Client Momentum
von: Gu, Xiaolan, et al.
Veröffentlicht: (2023)
von: Gu, Xiaolan, et al.
Veröffentlicht: (2023)
Efficient Differentially Private Fine-Tuning of Diffusion Models
von: Liu, Jing, et al.
Veröffentlicht: (2024)
von: Liu, Jing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning
von: Xia, Saisai, et al.
Veröffentlicht: (2025) -
TernaryVote: Differentially Private, Communication Efficient, and Byzantine Resilient Distributed Optimization on Heterogeneous Data
von: Jin, Richeng, et al.
Veröffentlicht: (2024) -
CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution
von: Zhu, Huifeng, et al.
Veröffentlicht: (2025) -
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
von: Xu, Tianshi, et al.
Veröffentlicht: (2024) -
UFO: Unlocking Ultra-Efficient Quantized Private Inference with Protocol and Algorithm Co-Optimization
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2026)