Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
Fuente:
arXiv
Salvato in:
| Autori principali: | Xie, Yuanbo, Zhang, Yingjie, Li, Yulin, Song, Shouyou, Chen, Xiaokun, Liu, Zhihan, Su, Liya, Liu, Tingwen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
di: Xie, Yuanbo, et al.
Pubblicazione: (2025)
di: Xie, Yuanbo, et al.
Pubblicazione: (2025)
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025)
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025)
SoK: Runtime Integrity
di: Ammar, Mahmoud, et al.
Pubblicazione: (2024)
di: Ammar, Mahmoud, et al.
Pubblicazione: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
di: Peng, Yuefeng, et al.
Pubblicazione: (2024)
di: Peng, Yuefeng, et al.
Pubblicazione: (2024)
Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems
di: Wang, Yihao, et al.
Pubblicazione: (2025)
di: Wang, Yihao, et al.
Pubblicazione: (2025)
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
di: Fan, Wei, et al.
Pubblicazione: (2024)
di: Fan, Wei, et al.
Pubblicazione: (2024)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
di: Liu, Yize, et al.
Pubblicazione: (2025)
di: Liu, Yize, et al.
Pubblicazione: (2025)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Runtime Backdoor Detection for Federated Learning via Representational Dissimilarity Analysis
di: Zhang, Xiyue, et al.
Pubblicazione: (2025)
di: Zhang, Xiyue, et al.
Pubblicazione: (2025)
ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking
di: Wu, Yutao, et al.
Pubblicazione: (2025)
di: Wu, Yutao, et al.
Pubblicazione: (2025)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
Privacy in Large Language Models: Attacks, Defenses and Future Directions
di: Li, Haoran, et al.
Pubblicazione: (2023)
di: Li, Haoran, et al.
Pubblicazione: (2023)
Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models
di: Gong, Yuyang, et al.
Pubblicazione: (2025)
di: Gong, Yuyang, et al.
Pubblicazione: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
di: Liu, Lijia, et al.
Pubblicazione: (2025)
di: Liu, Lijia, et al.
Pubblicazione: (2025)
Theorem-Carrying Transactions: Runtime Verification to Ensure Interface Specifications for Smart Contract Safety
di: Ball, Thomas, et al.
Pubblicazione: (2024)
di: Ball, Thomas, et al.
Pubblicazione: (2024)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2026)
di: Wang, Xilong, et al.
Pubblicazione: (2026)
Graded Modal Types for Integrity and Confidentiality
di: Marshall, Danielle, et al.
Pubblicazione: (2023)
di: Marshall, Danielle, et al.
Pubblicazione: (2023)
Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training
di: Li, Yuanfan, et al.
Pubblicazione: (2025)
di: Li, Yuanfan, et al.
Pubblicazione: (2025)
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
di: Shen, Guobin, et al.
Pubblicazione: (2024)
di: Shen, Guobin, et al.
Pubblicazione: (2024)
AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation
di: Liu, Zefang, et al.
Pubblicazione: (2025)
di: Liu, Zefang, et al.
Pubblicazione: (2025)
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models
di: Liu, Pang, et al.
Pubblicazione: (2026)
di: Liu, Pang, et al.
Pubblicazione: (2026)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
di: Hu, Xiao, et al.
Pubblicazione: (2025)
di: Hu, Xiao, et al.
Pubblicazione: (2025)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
di: Li, Haoran, et al.
Pubblicazione: (2024)
di: Li, Haoran, et al.
Pubblicazione: (2024)
ContraFix: Agentic Vulnerability Repair via Differential Runtime Evidence and Skill Reuse
di: Liu, Simiao, et al.
Pubblicazione: (2026)
di: Liu, Simiao, et al.
Pubblicazione: (2026)
Personalized Attacks of Social Engineering in Multi-turn Conversations: LLM Agents for Simulation and Detection
di: Kumarage, Tharindu, et al.
Pubblicazione: (2025)
di: Kumarage, Tharindu, et al.
Pubblicazione: (2025)
FORAY: Towards Effective Attack Synthesis against Deep Logical Vulnerabilities in DeFi Protocols
di: Wen, Hongbo, et al.
Pubblicazione: (2024)
di: Wen, Hongbo, et al.
Pubblicazione: (2024)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
di: Wang, Junlin, et al.
Pubblicazione: (2024)
di: Wang, Junlin, et al.
Pubblicazione: (2024)
Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
di: Min, Nay Myat, et al.
Pubblicazione: (2026)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
di: Zhang, Xiaozhe, et al.
Pubblicazione: (2026)
di: Zhang, Xiaozhe, et al.
Pubblicazione: (2026)
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
Valkyrie: A Response Framework to Augment Runtime Detection of Time-Progressive Attacks
di: Singh, Nikhilesh, et al.
Pubblicazione: (2025)
di: Singh, Nikhilesh, et al.
Pubblicazione: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
di: Fu, Wenjie, et al.
Pubblicazione: (2026)
di: Fu, Wenjie, et al.
Pubblicazione: (2026)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
di: An, Li, et al.
Pubblicazione: (2025)
di: An, Li, et al.
Pubblicazione: (2025)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
di: Huang, Chao, et al.
Pubblicazione: (2025)
di: Huang, Chao, et al.
Pubblicazione: (2025)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
di: More, Yash, et al.
Pubblicazione: (2024)
di: More, Yash, et al.
Pubblicazione: (2024)
Enabling Efficient Attack Investigation via Human-in-the-Loop Security Analysis
di: Tsegai, Saimon Amanuel, et al.
Pubblicazione: (2022)
di: Tsegai, Saimon Amanuel, et al.
Pubblicazione: (2022)
Michscan: Black-Box Neural Network Integrity Checking at Runtime Through Power Analysis
di: Paul, Robi, et al.
Pubblicazione: (2025)
di: Paul, Robi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
di: Xie, Yuanbo, et al.
Pubblicazione: (2025) -
Privacy Checklist: Privacy Violation Detection Grounding on Contextual Integrity Theory
di: Li, Haoran, et al.
Pubblicazione: (2024) -
DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack
di: Li, Hao, et al.
Pubblicazione: (2025) -
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
di: Zhao, Tianzhe, et al.
Pubblicazione: (2025) -
SoK: Runtime Integrity
di: Ammar, Mahmoud, et al.
Pubblicazione: (2024)