Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garimella, Karthik, Neda, Negar, Ebel, Austin, Jha, Nandan Kumar, Reagen, Brandon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HE-LRM: Efficient Private Embedding Lookups for Neural Inference Using Fully Homomorphic Encryption
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
DeepReShape: Redesigning Neural Networks for Efficient Private Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2023)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2023)
EinHops: Einsum Notation for Expressive Homomorphic Operations on RNS-CKKS Tensors
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
Orion: A Fully Homomorphic Encryption Framework for Deep Learning
von: Ebel, Austin, et al.
Veröffentlicht: (2023)
von: Ebel, Austin, et al.
Veröffentlicht: (2023)
AERO: Entropy-Guided Framework for Private LLM Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)
CiFlow: Dataflow Analysis and Optimization of Key Switching for Homomorphic Encryption
von: Neda, Negar, et al.
Veröffentlicht: (2023)
von: Neda, Negar, et al.
Veröffentlicht: (2023)
Entropy-Guided Attention for Private LLMs
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
TruncFormer: Private LLM Inference Using Only Truncations
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2024)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2024)
Osiris: A Systolic Approach to Accelerating Fully Homomorphic Encryption
von: Ebel, Austin, et al.
Veröffentlicht: (2024)
von: Ebel, Austin, et al.
Veröffentlicht: (2024)
NTTSuite: Number Theoretic Transform Benchmarks for Accelerating Encrypted Computation
von: Ding, Juran, et al.
Veröffentlicht: (2024)
von: Ding, Juran, et al.
Veröffentlicht: (2024)
SZKP: A Scalable Accelerator Architecture for Zero-Knowledge Proofs
von: Daftardar, Alhad, et al.
Veröffentlicht: (2024)
von: Daftardar, Alhad, et al.
Veröffentlicht: (2024)
CipherFormer: Efficient Transformer Private Inference with Low Round Complexity
von: Wang, Weize, et al.
Veröffentlicht: (2024)
von: Wang, Weize, et al.
Veröffentlicht: (2024)
CipherPrune: Efficient and Scalable Private Transformer Inference
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yancheng, et al.
Veröffentlicht: (2025)
CryptPEFT: Efficient and Private Neural Network Inference via Parameter-Efficient Fine-Tuning
von: Xia, Saisai, et al.
Veröffentlicht: (2025)
von: Xia, Saisai, et al.
Veröffentlicht: (2025)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2024)
Linearizing Models for Efficient yet Robust Private Inference
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
A Survey on Private Transformer Inference
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Hyena: Optimizing Homomorphically Encrypted Convolution for Private CNN Inference
von: Roh, Hyeri, et al.
Veröffentlicht: (2023)
von: Roh, Hyeri, et al.
Veröffentlicht: (2023)
Optimized Layerwise Approximation for Efficient Private Inference on Fully Homomorphic Encryption
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
PrivCirNet: Efficient Private Inference via Block Circulant Transformation
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC
von: Xu, Tianshi, et al.
Veröffentlicht: (2025)
von: Xu, Tianshi, et al.
Veröffentlicht: (2025)
Accelerating Private Large Transformers Inference through Fine-grained Collaborative Computation
von: Chen, Yuntian, et al.
Veröffentlicht: (2024)
von: Chen, Yuntian, et al.
Veröffentlicht: (2024)
Private Transformer Inference in MLaaS: A Survey
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Private, Efficient and Scalable Kernel Learning for Medical Image Analysis
von: Hannemann, Anika, et al.
Veröffentlicht: (2024)
von: Hannemann, Anika, et al.
Veröffentlicht: (2024)
Fast Homomorphic Linear Algebra with BLAS
von: Bae, Youngjin, et al.
Veröffentlicht: (2025)
von: Bae, Youngjin, et al.
Veröffentlicht: (2025)
Efficient and High-Accuracy Private CNN Inference with Helper-Assisted Malicious Security
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
UFO: Unlocking Ultra-Efficient Quantized Private Inference with Protocol and Algorithm Co-Optimization
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2026)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2026)
NeuJeans: Private Neural Network Inference with Joint Optimization of Convolution and FHE Bootstrapping
von: Ju, Jae Hyung, et al.
Veröffentlicht: (2023)
von: Ju, Jae Hyung, et al.
Veröffentlicht: (2023)
An Efficient Anomaly Detection Framework for Wireless Sensor Networks Using Markov Process
von: Mishra, Rahul, et al.
Veröffentlicht: (2025)
von: Mishra, Rahul, et al.
Veröffentlicht: (2025)
MOFHEI: Model Optimizing Framework for Fast and Efficient Homomorphically Encrypted Neural Network Inference
von: Ghazvinian, Parsa, et al.
Veröffentlicht: (2024)
von: Ghazvinian, Parsa, et al.
Veröffentlicht: (2024)
zkPHIRE: A Programmable Accelerator for ZKPs over HIgh-degRee, Expressive Gates
von: Daftardar, Alhad, et al.
Veröffentlicht: (2025)
von: Daftardar, Alhad, et al.
Veröffentlicht: (2025)
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
Differentially Private Linear Optimization for Multi-Party Resource Sharing
von: Karaca, Utku, et al.
Veröffentlicht: (2021)
von: Karaca, Utku, et al.
Veröffentlicht: (2021)
ShadowBound: Efficient Heap Memory Protection Through Advanced Metadata Management and Customized Compiler Optimization
von: Yu, Zheng, et al.
Veröffentlicht: (2024)
von: Yu, Zheng, et al.
Veröffentlicht: (2024)
Disparate Impact on Group Accuracy of Linearization for Private Inference
von: Das, Saswat, et al.
Veröffentlicht: (2024)
von: Das, Saswat, et al.
Veröffentlicht: (2024)
NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
Differentially Private Inference for Longitudinal Linear Regression
von: Sopa, Getoar, et al.
Veröffentlicht: (2026)
von: Sopa, Getoar, et al.
Veröffentlicht: (2026)
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HE-LRM: Efficient Private Embedding Lookups for Neural Inference Using Fully Homomorphic Encryption
von: Garimella, Karthik, et al.
Veröffentlicht: (2025) -
DeepReShape: Redesigning Neural Networks for Efficient Private Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2023) -
EinHops: Einsum Notation for Expressive Homomorphic Operations on RNS-CKKS Tensors
von: Garimella, Karthik, et al.
Veröffentlicht: (2025) -
Orion: A Fully Homomorphic Encryption Framework for Deep Learning
von: Ebel, Austin, et al.
Veröffentlicht: (2023) -
AERO: Entropy-Guided Framework for Private LLM Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)