TruncFormer: Private LLM Inference Using Only Truncations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yubeaton, Patrick, Mo, Jianqiao Cambridge, Garimella, Karthik, Jha, Nandan Kumar, Reagen, Brandon, Hegde, Chinmay, Garg, Siddharth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AERO: Entropy-Guided Framework for Private LLM Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)
Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
DeepReShape: Redesigning Neural Networks for Efficient Private Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2023)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2023)
Entropy-Guided Attention for Private LLMs
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
HE-LRM: Efficient Private Embedding Lookups for Neural Inference Using Fully Homomorphic Encryption
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
EinHops: Einsum Notation for Expressive Homomorphic Operations on RNS-CKKS Tensors
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)
Orion: A Fully Homomorphic Encryption Framework for Deep Learning
von: Ebel, Austin, et al.
Veröffentlicht: (2023)
von: Ebel, Austin, et al.
Veröffentlicht: (2023)
Exploring the Agentic Frontier of Verilog Code Generation
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2026)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2026)
zkPHIRE: A Programmable Accelerator for ZKPs over HIgh-degRee, Expressive Gates
von: Daftardar, Alhad, et al.
Veröffentlicht: (2025)
von: Daftardar, Alhad, et al.
Veröffentlicht: (2025)
SZKP: A Scalable Accelerator Architecture for Zero-Knowledge Proofs
von: Daftardar, Alhad, et al.
Veröffentlicht: (2024)
von: Daftardar, Alhad, et al.
Veröffentlicht: (2024)
Need for zkSpeed: Accelerating HyperPlonk for Zero-Knowledge Proofs
von: Daftardar, Alhad, et al.
Veröffentlicht: (2025)
von: Daftardar, Alhad, et al.
Veröffentlicht: (2025)
NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward Networks
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space?
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
A Random Matrix Theory Perspective on the Learning Dynamics of Multi-head Latent Attention
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025)
ReLU's Revival: On the Entropic Overload in Normalization-Free Large Language Models
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024)
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
Osiris: A Systolic Approach to Accelerating Fully Homomorphic Encryption
von: Ebel, Austin, et al.
Veröffentlicht: (2024)
von: Ebel, Austin, et al.
Veröffentlicht: (2024)
Huff-LLM: End-to-End Lossless Compression for Efficient LLM Inference
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
CipherFormer: Efficient Transformer Private Inference with Low Round Complexity
von: Wang, Weize, et al.
Veröffentlicht: (2024)
von: Wang, Weize, et al.
Veröffentlicht: (2024)
VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
NTTSuite: Number Theoretic Transform Benchmarks for Accelerating Encrypted Computation
von: Ding, Juran, et al.
Veröffentlicht: (2024)
von: Ding, Juran, et al.
Veröffentlicht: (2024)
CiFlow: Dataflow Analysis and Optimization of Key Switching for Homomorphic Encryption
von: Neda, Negar, et al.
Veröffentlicht: (2023)
von: Neda, Negar, et al.
Veröffentlicht: (2023)
Discovering Sparse Recovery Algorithms Using Neural Architecture Search
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
Differentially Private Data Generation with Missing Data
von: Mohapatra, Shubhankar, et al.
Veröffentlicht: (2023)
von: Mohapatra, Shubhankar, et al.
Veröffentlicht: (2023)
GOTCHA: Real-Time Video Deepfake Detection via Challenge-Response
von: Mittal, Govind, et al.
Veröffentlicht: (2022)
von: Mittal, Govind, et al.
Veröffentlicht: (2022)
MTU: The Multifunction Tree Unit for Accelerating Zero-Knowledge Proofs
von: Mo, Jianqiao, et al.
Veröffentlicht: (2025)
von: Mo, Jianqiao, et al.
Veröffentlicht: (2025)
Private Statistical Estimation via Truncation
von: Zampetakis, Manolis, et al.
Veröffentlicht: (2025)
von: Zampetakis, Manolis, et al.
Veröffentlicht: (2025)
Cascade: Token-Sharded Private LLM Inference
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)
SelectFormer: Private and Practical Data Selection for Transformers
von: Ouyang, Xu, et al.
Veröffentlicht: (2023)
von: Ouyang, Xu, et al.
Veröffentlicht: (2023)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2025)
PermLLM: Private Inference of Large Language Models within 3 Seconds under WAN
von: Zheng, Fei, et al.
Veröffentlicht: (2024)
von: Zheng, Fei, et al.
Veröffentlicht: (2024)
SEAL: Semantic Aware Image Watermarking
von: Arabi, Kasra, et al.
Veröffentlicht: (2025)
von: Arabi, Kasra, et al.
Veröffentlicht: (2025)
FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
von: Lin, Chenqi, et al.
Veröffentlicht: (2024)
von: Lin, Chenqi, et al.
Veröffentlicht: (2024)
An LLM Framework For Cryptography Over Chat Channels
von: Gligoroski, Danilo, et al.
Veröffentlicht: (2025)
von: Gligoroski, Danilo, et al.
Veröffentlicht: (2025)
Towards Secure and Private AI: A Framework for Decentralized Inference
von: Zhang, Hongyang, et al.
Veröffentlicht: (2024)
von: Zhang, Hongyang, et al.
Veröffentlicht: (2024)
PITCH: AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response
von: Mittal, Govind, et al.
Veröffentlicht: (2024)
von: Mittal, Govind, et al.
Veröffentlicht: (2024)
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification
von: Feuer, Benjamin, et al.
Veröffentlicht: (2024)
von: Feuer, Benjamin, et al.
Veröffentlicht: (2024)
Surgical Repair of Insecure Code Generation in LLMs
von: Sandoval, Gustavo, et al.
Veröffentlicht: (2026)
von: Sandoval, Gustavo, et al.
Veröffentlicht: (2026)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
von: Wu, Fangzhou, et al.
Veröffentlicht: (2024)
A Survey on Private Transformer Inference
von: Li, Yang, et al.
Veröffentlicht: (2024)
von: Li, Yang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AERO: Entropy-Guided Framework for Private LLM Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2024) -
Network and Compiler Optimizations for Efficient Linear Algebra Kernels in Private Transformer Inference
von: Garimella, Karthik, et al.
Veröffentlicht: (2025) -
DeepReShape: Redesigning Neural Networks for Efficient Private Inference
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2023) -
Entropy-Guided Attention for Private LLMs
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2025) -
HE-LRM: Efficient Private Embedding Lookups for Neural Inference Using Fully Homomorphic Encryption
von: Garimella, Karthik, et al.
Veröffentlicht: (2025)