Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuhao, Qu, Wenjie, Zhai, Shengfang, Jiang, Yanze, Liu, Zichen, Liu, Yue, Dong, Yinpeng, Zhang, Jiaheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026)
by: Wang, Yuhao, et al.
Published: (2026)
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
by: Zhai, Shengfang, et al.
Published: (2025)
by: Zhai, Shengfang, et al.
Published: (2025)
Securing LLM Agents Need Intent-to-Execution Integrity
by: Qu, Wenjie, et al.
Published: (2026)
by: Qu, Wenjie, et al.
Published: (2026)
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
by: Zhai, Shengfang, et al.
Published: (2026)
by: Zhai, Shengfang, et al.
Published: (2026)
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
by: Wu, Linyu, et al.
Published: (2025)
by: Wu, Linyu, et al.
Published: (2025)
Discovering Universal Semantic Triggers for Text-to-Image Synthesis
by: Zhai, Shengfang, et al.
Published: (2024)
by: Zhai, Shengfang, et al.
Published: (2024)
IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
by: Guo, Yanpei, et al.
Published: (2026)
by: Guo, Yanpei, et al.
Published: (2026)
Provably Robust Multi-bit Watermarking for AI-generated Text
by: Qu, Wenjie, et al.
Published: (2024)
by: Qu, Wenjie, et al.
Published: (2024)
Silent Egress: When Implicit Prompt Injection Makes LLM Agents Leak Without a Trace
by: Lan, Qianlong, et al.
Published: (2026)
by: Lan, Qianlong, et al.
Published: (2026)
Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models
by: Tian, Zhihua, et al.
Published: (2025)
by: Tian, Zhihua, et al.
Published: (2025)
ExtendAttack: Attacking Servers of LRMs via Extending Reasoning
by: Zhu, Zhenhao, et al.
Published: (2025)
by: Zhu, Zhenhao, et al.
Published: (2025)
BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
S-Leak: Leakage-Abuse Attack Against Efficient Conjunctive SSE via s-term Leakage
by: Su, Yue, et al.
Published: (2025)
by: Su, Yue, et al.
Published: (2025)
ThreatPilot: Attack-Driven Threat Intelligence Extraction
by: Xu, Ming, et al.
Published: (2024)
by: Xu, Ming, et al.
Published: (2024)
The Silent Spill: Measuring Sensitive Data Leaks Across Public URL Repositories
by: Ramadan, Tarek, et al.
Published: (2026)
by: Ramadan, Tarek, et al.
Published: (2026)
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy
by: Zhai, Shengfang, et al.
Published: (2024)
by: Zhai, Shengfang, et al.
Published: (2024)
Leaking Queries On Secure Stream Processing Systems
by: Pham, Hung, et al.
Published: (2025)
by: Pham, Hung, et al.
Published: (2025)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation
by: Qi, Zhisheng, et al.
Published: (2026)
by: Qi, Zhisheng, et al.
Published: (2026)
Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
Five Queries Are Enough: Query-Efficient and Surrogate-Free Membership Inference Attacks on RAG via Entailment
by: Nguyen, Nguyen Linh Bao, et al.
Published: (2026)
by: Nguyen, Nguyen Linh Bao, et al.
Published: (2026)
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
by: Liu, Tong, et al.
Published: (2024)
by: Liu, Tong, et al.
Published: (2024)
Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
by: Xue, Jing, et al.
Published: (2025)
by: Xue, Jing, et al.
Published: (2025)
Coordinated Position Falsification Attacks and Countermeasures for Location-Based Services
by: Liu, Wenjie, et al.
Published: (2025)
by: Liu, Wenjie, et al.
Published: (2025)
LeakDojo: Decoding the Leakage Threats of RAG Systems
by: Zhang, Maosen, et al.
Published: (2026)
by: Zhang, Maosen, et al.
Published: (2026)
Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity
by: Guo, Wei, et al.
Published: (2025)
by: Guo, Wei, et al.
Published: (2025)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
by: Zhao, Tianzhe, et al.
Published: (2025)
by: Zhao, Tianzhe, et al.
Published: (2025)
Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks
by: Al-Lawati, Ali, et al.
Published: (2026)
by: Al-Lawati, Ali, et al.
Published: (2026)
Effectiveness of Adversarial Benign and Malware Examples in Evasion and Poisoning Attacks
by: Kozák, Matouš, et al.
Published: (2025)
by: Kozák, Matouš, et al.
Published: (2025)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
by: Kaneko, Masahiro, et al.
Published: (2025)
by: Kaneko, Masahiro, et al.
Published: (2025)
Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
by: Xie, Yuanbo, et al.
Published: (2026)
by: Xie, Yuanbo, et al.
Published: (2026)
EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling
by: Song, Zichen, et al.
Published: (2024)
by: Song, Zichen, et al.
Published: (2024)
ATOM: A Framework of Detecting Query-Based Model Extraction Attacks for Graph Neural Networks
by: Cheng, Zhan, et al.
Published: (2025)
by: Cheng, Zhan, et al.
Published: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
"Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval
by: Li, Jiate, et al.
Published: (2026)
by: Li, Jiate, et al.
Published: (2026)
Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
by: Xie, Zhixin, et al.
Published: (2025)
by: Xie, Zhixin, et al.
Published: (2025)
A Benign Activity Extraction Method for Malignant Activity Identification using Data Provenance
by: Saito, Taishin
Published: (2025)
by: Saito, Taishin
Published: (2025)
Similar Items
-
MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots
by: Wang, Yuhao, et al.
Published: (2026) -
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
by: Zhai, Shengfang, et al.
Published: (2025) -
Securing LLM Agents Need Intent-to-Execution Integrity
by: Qu, Wenjie, et al.
Published: (2026) -
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
by: Zhai, Shengfang, et al.
Published: (2026) -
DMark: Order-Agnostic Watermarking for Diffusion Large Language Models
by: Wu, Linyu, et al.
Published: (2025)