Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
Fuente:
arXiv
Saved in:
| Main Authors: | Chrapek, Marcin, Copik, Marcin, Mettaz, Etienne, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
Confidential Computing on Heterogeneous CPU-GPU Systems: Survey and Future Directions
by: Wang, Qifan, et al.
Published: (2024)
by: Wang, Qifan, et al.
Published: (2024)
Fastrack: Fast IO for Secure ML using GPU TEEs
by: Wang, Yongqin, et al.
Published: (2024)
by: Wang, Yongqin, et al.
Published: (2024)
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026)
by: Chau, Vu Minh, et al.
Published: (2026)
CiFlow: Dataflow Analysis and Optimization of Key Switching for Homomorphic Encryption
by: Neda, Negar, et al.
Published: (2023)
by: Neda, Negar, et al.
Published: (2023)
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
by: Nasr-Esfahany, Arash, et al.
Published: (2025)
GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference
by: Chen, Guoci, et al.
Published: (2026)
by: Chen, Guoci, et al.
Published: (2026)
Revealing CNN Architectures via Side-Channel Analysis in Dataflow-based Inference Accelerators
by: Weerasena, Hansika, et al.
Published: (2023)
by: Weerasena, Hansika, et al.
Published: (2023)
The Impact of Logic Locking on Confidentiality: An Automated Evaluation
by: Reimann, Lennart M., et al.
Published: (2025)
by: Reimann, Lennart M., et al.
Published: (2025)
Research Directions for Verifiable Crypto-Physically Secure TEEs
by: Bellemare, Sylvain
Published: (2024)
by: Bellemare, Sylvain
Published: (2024)
VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems
by: Yamamoto, Yuji, et al.
Published: (2026)
by: Yamamoto, Yuji, et al.
Published: (2026)
The Road to Trust: Building Enclaves within Confidential VMs
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
VeriLeaky: Navigating IP Protection vs Utility in Fine-Tuning for LLM-Driven Verilog Coding
by: Wang, Zeng, et al.
Published: (2025)
by: Wang, Zeng, et al.
Published: (2025)
An Early Experience with Confidential Computing Architecture for On-Device Model Protection
by: Abdollahi, Sina, et al.
Published: (2025)
by: Abdollahi, Sina, et al.
Published: (2025)
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs
by: Verma, Tarunesh, et al.
Published: (2025)
by: Verma, Tarunesh, et al.
Published: (2025)
Side-channel Inference of User Activities in AR/VR Using GPU Profiling
by: Son, Seonghun, et al.
Published: (2025)
by: Son, Seonghun, et al.
Published: (2025)
Hawkeye: Reproducing GPU-Level Non-Determinism
by: Badash, Erez, et al.
Published: (2026)
by: Badash, Erez, et al.
Published: (2026)
KeyVisor -- A Lightweight ISA Extension for Protected Key Handles with CPU-enforced Usage Policies
by: Schwarz, Fabian, et al.
Published: (2024)
by: Schwarz, Fabian, et al.
Published: (2024)
WebGPU-SPY: Finding Fingerprints in the Sandbox through GPU Cache Attacks
by: Ferguson, Ethan, et al.
Published: (2024)
by: Ferguson, Ethan, et al.
Published: (2024)
GCL-Sampler: Discovering Kernel Similarity for Sampled GPU Simulation via Graph Contrastive Learning
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
FedBit: Accelerating Privacy-Preserving Federated Learning via Bit-Interleaved Packing and Cross-Layer Co-Design
by: Meng, Xiangchen, et al.
Published: (2025)
by: Meng, Xiangchen, et al.
Published: (2025)
Attack on a PUF-based Secure Binary Neural Network
by: Basak, Bijeet, et al.
Published: (2025)
by: Basak, Bijeet, et al.
Published: (2025)
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning
by: Tol, M. Caner, et al.
Published: (2025)
by: Tol, M. Caner, et al.
Published: (2025)
SafeLight: Enhancing Security in Optical Convolutional Neural Network Accelerators
by: Afifi, Salma, et al.
Published: (2024)
by: Afifi, Salma, et al.
Published: (2024)
LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges
by: Knechtel, Johann, et al.
Published: (2026)
by: Knechtel, Johann, et al.
Published: (2026)
DL2Fence: Integrating Deep Learning and Frame Fusion for Enhanced Detection and Localization of Refined Denial-of-Service in Large-Scale NoCs
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Tessera: Secure, Near-Line-Rate Weight Streaming for UMA Edge Accelerators
by: Naskar, Animan
Published: (2026)
by: Naskar, Animan
Published: (2026)
Breaking On-Chip Communication Anonymity using Flow Correlation Attacks
by: Weerasena, Hansika, et al.
Published: (2023)
by: Weerasena, Hansika, et al.
Published: (2023)
Evasive Hardware Trojan through Adversarial Power Trace
by: Omidi, Behnam, et al.
Published: (2024)
by: Omidi, Behnam, et al.
Published: (2024)
2-in-1 Accelerator: Enabling Random Precision Switch for Winning Both Adversarial Robustness and Efficiency
by: Fu, Yonggan, et al.
Published: (2021)
by: Fu, Yonggan, et al.
Published: (2021)
IMS: Intelligent Hardware Monitoring System for Secure SoCs
by: Foudhaili, Wadid, et al.
Published: (2026)
by: Foudhaili, Wadid, et al.
Published: (2026)
Thales: Formulating and Estimating Architectural Vulnerability Factors for DNN Accelerators
by: Tyagi, Abhishek, et al.
Published: (2022)
by: Tyagi, Abhishek, et al.
Published: (2022)
TroLLoc: Logic Locking and Layout Hardening for IC Security Closure against Hardware Trojans
by: Wang, Fangzhou, et al.
Published: (2024)
by: Wang, Fangzhou, et al.
Published: (2024)
LLMs and the Future of Chip Design: Unveiling Security Risks and Building Trust
by: Wang, Zeng, et al.
Published: (2024)
by: Wang, Zeng, et al.
Published: (2024)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
by: Patwari, Rajeev, et al.
Published: (2025)
by: Patwari, Rajeev, et al.
Published: (2025)
Vulnerabilities in Partial TEE-Shielded LLM Inference with Precomputed Noise
by: Saini, Abhishek, et al.
Published: (2026)
by: Saini, Abhishek, et al.
Published: (2026)
Cost-Effective Optimization and Implementation of the CRT-Paillier Decryption Algorithm for Enhanced Performance
by: Huang, Zhengwu, et al.
Published: (2025)
by: Huang, Zhengwu, et al.
Published: (2025)
GPU in the Blind Spot: Overlooked Security Risks in Transportation
by: Puspa, Sefatun-Noor, et al.
Published: (2025)
by: Puspa, Sefatun-Noor, et al.
Published: (2025)
FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption
by: Daksha, Lohit, et al.
Published: (2026)
by: Daksha, Lohit, et al.
Published: (2026)
Similar Items
-
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025) -
Confidential Computing on Heterogeneous CPU-GPU Systems: Survey and Future Directions
by: Wang, Qifan, et al.
Published: (2024) -
Fastrack: Fast IO for Secure ML using GPU TEEs
by: Wang, Yongqin, et al.
Published: (2024) -
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026) -
CiFlow: Dataflow Analysis and Optimization of Key Switching for Homomorphic Encryption
by: Neda, Negar, et al.
Published: (2023)