TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Pengzhou, Ding, Yidong, Ju, Tianjie, Wu, Zongru, Du, Wei, Yi, Ping, Zhang, Zhuosheng, Liu, Gongshen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transferring Backdoors between Large Language Models by Knowledge Distillation
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
von: Wu, Zongru, et al.
Veröffentlicht: (2024)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
MKF-ADS: Multi-Knowledge Fusion Based Self-supervised Anomaly Detection System for Control Area Network
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data
von: Zhao, Haodong, et al.
Veröffentlicht: (2026)
von: Zhao, Haodong, et al.
Veröffentlicht: (2026)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
von: Zhao, Tianhang, et al.
Veröffentlicht: (2025)
von: Zhao, Tianhang, et al.
Veröffentlicht: (2025)
CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
von: Zhao, Haodong, et al.
Veröffentlicht: (2024)
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
von: Clop, Cody, et al.
Veröffentlicht: (2024)
von: Clop, Cody, et al.
Veröffentlicht: (2024)
Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2025)
Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective
von: Zhao, Haodong, et al.
Veröffentlicht: (2026)
von: Zhao, Haodong, et al.
Veröffentlicht: (2026)
Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks
von: Bagwe, Gaurav, et al.
Veröffentlicht: (2025)
von: Bagwe, Gaurav, et al.
Veröffentlicht: (2025)
LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
von: Yan, Zihe, et al.
Veröffentlicht: (2025)
von: Yan, Zihe, et al.
Veröffentlicht: (2025)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
von: Du, Wei, et al.
Veröffentlicht: (2023)
von: Du, Wei, et al.
Veröffentlicht: (2023)
EmbTracker: Traceable Black-box Watermarking for Federated Language Models
von: Zhao, Haodong, et al.
Veröffentlicht: (2026)
von: Zhao, Haodong, et al.
Veröffentlicht: (2026)
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
von: Zou, Wei, et al.
Veröffentlicht: (2024)
von: Zou, Wei, et al.
Veröffentlicht: (2024)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
von: Wu, Zongru, et al.
Veröffentlicht: (2025)
SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
von: Feng, Yunhao, et al.
Veröffentlicht: (2026)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models
von: Zhang, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yucheng, et al.
Veröffentlicht: (2024)
Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
von: Li, Tian, et al.
Veröffentlicht: (2025)
von: Li, Tian, et al.
Veröffentlicht: (2025)
SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
von: Lv, Peizhuo, et al.
Veröffentlicht: (2025)
A Universal Identity Backdoor Attack against Speaker Verification based on Siamese Network
von: Zhao, Haodong, et al.
Veröffentlicht: (2023)
von: Zhao, Haodong, et al.
Veröffentlicht: (2023)
The Philosopher's Stone: Trojaning Plugins of Large Language Models
von: Dong, Tian, et al.
Veröffentlicht: (2023)
von: Dong, Tian, et al.
Veröffentlicht: (2023)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data
von: Zeng, Shenglai, et al.
Veröffentlicht: (2024)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2024)
Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
TechniqueRAG: Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
von: Lekssays, Ahmed, et al.
Veröffentlicht: (2025)
von: Lekssays, Ahmed, et al.
Veröffentlicht: (2025)
Evil from Within: Machine Learning Backdoors through Hardware Trojans
von: Warnecke, Alexander, et al.
Veröffentlicht: (2023)
von: Warnecke, Alexander, et al.
Veröffentlicht: (2023)
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Tianzhe, et al.
Veröffentlicht: (2025)
Event Trojan: Asynchronous Event-based Backdoor Attacks
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
von: Wang, Ruofei, et al.
Veröffentlicht: (2024)
Neural Trojans
von: Liu, Yuntao, et al.
Veröffentlicht: (2017)
von: Liu, Yuntao, et al.
Veröffentlicht: (2017)
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2024)
von: Zeng, Shenglai, et al.
Veröffentlicht: (2024)
LATENT: LLM-Augmented Trojan Insertion and Evaluation Framework for Analog Netlist Topologies
von: Chaudhuri, Jayeeta, et al.
Veröffentlicht: (2025)
von: Chaudhuri, Jayeeta, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transferring Backdoors between Large Language Models by Knowledge Distillation
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024) -
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
von: Wu, Zongru, et al.
Veröffentlicht: (2024) -
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
von: Wu, Zongru, et al.
Veröffentlicht: (2024) -
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2023) -
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)