FastQuery: Communication-efficient Embedding Table Query for Private LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Chenqi, Xu, Tianshi, Yang, Zebin, Wang, Runsheng, Huang, Ru, Li, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
PrivCirNet: Efficient Private Inference via Block Circulant Transformation
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC
by: Xu, Tianshi, et al.
Published: (2025)
by: Xu, Tianshi, et al.
Published: (2025)
QueryCheetah: Fast Automated Discovery of Attribute Inference Attacks Against Query-Based Systems
by: Stevanoski, Bozhidar, et al.
Published: (2024)
by: Stevanoski, Bozhidar, et al.
Published: (2024)
UFO: Unlocking Ultra-Efficient Quantized Private Inference with Protocol and Algorithm Co-Optimization
by: Zeng, Wenxuan, et al.
Published: (2026)
by: Zeng, Wenxuan, et al.
Published: (2026)
EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
by: Zeng, Wenxuan, et al.
Published: (2024)
by: Zeng, Wenxuan, et al.
Published: (2024)
Membership Inference for Contrastive Pre-training Models with Text-only PII Queries
by: Cheng, Ruoxi, et al.
Published: (2026)
by: Cheng, Ruoxi, et al.
Published: (2026)
Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference
by: Xu, Xiangrui, et al.
Published: (2024)
by: Xu, Xiangrui, et al.
Published: (2024)
QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents
by: Xie, Yuchong, et al.
Published: (2025)
by: Xie, Yuchong, et al.
Published: (2025)
Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
A Survey on Private Transformer Inference
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Private Transformer Inference in MLaaS: A Survey
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language
by: Zou, Qingsong, et al.
Published: (2025)
by: Zou, Qingsong, et al.
Published: (2025)
QUEEN: Query Unlearning against Model Extraction
by: Chen, Huajie, et al.
Published: (2024)
by: Chen, Huajie, et al.
Published: (2024)
Differentially Private and Communication Efficient Large Language Model Split Inference via Stochastic Quantization and Soft Prompt
by: Gu, Yujie, et al.
Published: (2026)
by: Gu, Yujie, et al.
Published: (2026)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying
by: Lyu, Xingyu, et al.
Published: (2026)
by: Lyu, Xingyu, et al.
Published: (2026)
Fluent: Round-efficient Secure Aggregation for Private Federated Learning
by: Li, Xincheng, et al.
Published: (2024)
by: Li, Xincheng, et al.
Published: (2024)
Not My Agent, Not My Boundary? Elicitation of Personal Privacy Boundaries in AI-Delegated Information Sharing
by: Guo, Bingcan, et al.
Published: (2025)
by: Guo, Bingcan, et al.
Published: (2025)
Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset
by: Li, Tianshi
Published: (2026)
by: Li, Tianshi
Published: (2026)
Comet: Accelerating Private Inference for Large Language Model by Predicting Activation Sparsity
by: Yan, Guang, et al.
Published: (2025)
by: Yan, Guang, et al.
Published: (2025)
Exposing Hidden Interfaces: LLM-Guided Type Inference for Reverse Engineering macOS Private Frameworks
by: Kharlamova, Arina, et al.
Published: (2026)
by: Kharlamova, Arina, et al.
Published: (2026)
Towards Small Language Models for Security Query Generation in SOC Workflows
by: Muzammil, Saleha, et al.
Published: (2025)
by: Muzammil, Saleha, et al.
Published: (2025)
Towards Secure and Private AI: A Framework for Decentralized Inference
by: Zhang, Hongyang, et al.
Published: (2024)
by: Zhang, Hongyang, et al.
Published: (2024)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
by: Chen, Wenyu, et al.
Published: (2026)
by: Chen, Wenyu, et al.
Published: (2026)
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots
by: Zhou, Jijie, et al.
Published: (2024)
by: Zhou, Jijie, et al.
Published: (2024)
Exploring Membership Inference Vulnerabilities in Clinical Large Language Models
by: Nemecek, Alexander, et al.
Published: (2025)
by: Nemecek, Alexander, et al.
Published: (2025)
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
An Investigation on Group Query Hallucination Attacks
by: Miao, Kehao, et al.
Published: (2025)
by: Miao, Kehao, et al.
Published: (2025)
Towards Efficient Privacy-Preserving Machine Learning: A Systematic Review from Protocol, Model, and System Perspectives
by: Zeng, Wenxuan, et al.
Published: (2025)
by: Zeng, Wenxuan, et al.
Published: (2025)
Optimized Layerwise Approximation for Efficient Private Inference on Fully Homomorphic Encryption
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
SynRAG: A Large Language Model Framework for Executable Query Generation in Heterogeneous SIEM System
by: Saju, Md Hasan, et al.
Published: (2025)
by: Saju, Md Hasan, et al.
Published: (2025)
Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating Intent
by: Shang, Shang, et al.
Published: (2024)
by: Shang, Shang, et al.
Published: (2024)
FastFHE: Packing-Scalable and Depthwise-Separable CNN Inference Over FHE
by: Song, Wenbo, et al.
Published: (2025)
by: Song, Wenbo, et al.
Published: (2025)
MPCache: MPC-Friendly KV Cache Eviction for Efficient Private LLM Inference
by: Zeng, Wenxuan, et al.
Published: (2025)
by: Zeng, Wenxuan, et al.
Published: (2025)
Exploring Query Efficient Data Generation towards Data-free Model Stealing in Hard Label Setting
by: Pei, Gaozheng, et al.
Published: (2024)
by: Pei, Gaozheng, et al.
Published: (2024)
Differentially Private Distance Query with Asymmetric Noise
by: Sheng, Weihong, et al.
Published: (2025)
by: Sheng, Weihong, et al.
Published: (2025)
A First Look At Efficient And Secure On-Device LLM Inference Against KV Leakage
by: Yang, Huan, et al.
Published: (2024)
by: Yang, Huan, et al.
Published: (2024)
Similar Items
-
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
by: Xu, Tianshi, et al.
Published: (2024) -
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
by: Xu, Tianshi, et al.
Published: (2024) -
PrivCirNet: Efficient Private Inference via Block Circulant Transformation
by: Xu, Tianshi, et al.
Published: (2024) -
Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC
by: Xu, Tianshi, et al.
Published: (2025) -
QueryCheetah: Fast Automated Discovery of Attribute Inference Attacks Against Query-Based Systems
by: Stevanoski, Bozhidar, et al.
Published: (2024)