Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Zongru, Cheng, Pengzhou, Fang, Lingyong, Zhang, Zhuosheng, Liu, Gongshen |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
par: Wu, Zongru, et autres
Publié: (2024)
par: Wu, Zongru, et autres
Publié: (2024)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
par: Cheng, Pengzhou, et autres
Publié: (2024)
par: Cheng, Pengzhou, et autres
Publié: (2024)
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
par: Cheng, Pengzhou, et autres
Publié: (2024)
par: Cheng, Pengzhou, et autres
Publié: (2024)
MKF-ADS: Multi-Knowledge Fusion Based Self-supervised Anomaly Detection System for Control Area Network
par: Cheng, Pengzhou, et autres
Publié: (2024)
par: Cheng, Pengzhou, et autres
Publié: (2024)
Transferring Backdoors between Large Language Models by Knowledge Distillation
par: Cheng, Pengzhou, et autres
Publié: (2024)
par: Cheng, Pengzhou, et autres
Publié: (2024)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
par: Cheng, Pengzhou, et autres
Publié: (2023)
par: Cheng, Pengzhou, et autres
Publié: (2023)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
par: Du, Wei, et autres
Publié: (2023)
par: Du, Wei, et autres
Publié: (2023)
NSmark: Null Space Based Black-box Watermarking Defense Framework for Language Models
par: Zhao, Haodong, et autres
Publié: (2024)
par: Zhao, Haodong, et autres
Publié: (2024)
LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
par: Yan, Zihe, et autres
Publié: (2025)
par: Yan, Zihe, et autres
Publié: (2025)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
par: Zhao, Shuai, et autres
Publié: (2024)
par: Zhao, Shuai, et autres
Publié: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
par: Zhao, Shuai, et autres
Publié: (2024)
par: Zhao, Shuai, et autres
Publié: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
par: Cao, Yuanpu, et autres
Publié: (2023)
par: Cao, Yuanpu, et autres
Publié: (2023)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
par: Jiang, Peihai, et autres
Publié: (2025)
par: Jiang, Peihai, et autres
Publié: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
par: Xu, Jiashu, et autres
Publié: (2023)
par: Xu, Jiashu, et autres
Publié: (2023)
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models
par: Chen, Zhuo, et autres
Publié: (2024)
par: Chen, Zhuo, et autres
Publié: (2024)
ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data
par: Zhao, Haodong, et autres
Publié: (2026)
par: Zhao, Haodong, et autres
Publié: (2026)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
par: Hu, Man, et autres
Publié: (2025)
par: Hu, Man, et autres
Publié: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
par: Li, Xi, et autres
Publié: (2024)
par: Li, Xi, et autres
Publié: (2024)
ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
par: Cheng, Siyang, et autres
Publié: (2025)
par: Cheng, Siyang, et autres
Publié: (2025)
Goal-guided Generative Prompt Injection Attack on Large Language Models
par: Zhang, Chong, et autres
Publié: (2024)
par: Zhang, Chong, et autres
Publié: (2024)
Backdooring Bias in Large Language Models
par: Das, Anudeep, et autres
Publié: (2026)
par: Das, Anudeep, et autres
Publié: (2026)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
par: Ji, Wence, et autres
Publié: (2025)
par: Ji, Wence, et autres
Publié: (2025)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
par: Wu, Zongru, et autres
Publié: (2025)
par: Wu, Zongru, et autres
Publié: (2025)
Exploring Backdoor Vulnerabilities of Chat Models
par: Hao, Yunzhuo, et autres
Publié: (2024)
par: Hao, Yunzhuo, et autres
Publié: (2024)
REEF: Representation Encoding Fingerprints for Large Language Models
par: Zhang, Jie, et autres
Publié: (2024)
par: Zhang, Jie, et autres
Publié: (2024)
CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
par: Li, Yuetai, et autres
Publié: (2024)
par: Li, Yuetai, et autres
Publié: (2024)
Resource Consumption Threats in Large Language Models
par: Zhang, Yuanhe, et autres
Publié: (2026)
par: Zhang, Yuanhe, et autres
Publié: (2026)
Lightweight and Fast Backdoor Model Detection
par: Yu, Yinbo, et autres
Publié: (2026)
par: Yu, Yinbo, et autres
Publié: (2026)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
par: Liu, Qin, et autres
Publié: (2024)
par: Liu, Qin, et autres
Publié: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
par: Wei, Jiali, et autres
Publié: (2026)
par: Wei, Jiali, et autres
Publié: (2026)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
par: Liu, Junhao, et autres
Publié: (2025)
par: Liu, Junhao, et autres
Publié: (2025)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
par: Zhang, Jinchuan, et autres
Publié: (2024)
par: Zhang, Jinchuan, et autres
Publié: (2024)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
par: Yu, Miao, et autres
Publié: (2024)
par: Yu, Miao, et autres
Publié: (2024)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
par: Liu, Xuxu, et autres
Publié: (2025)
par: Liu, Xuxu, et autres
Publié: (2025)
Internal Safety Collapse in Frontier Large Language Models
par: Wu, Yutao, et autres
Publié: (2026)
par: Wu, Yutao, et autres
Publié: (2026)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
par: Du, Yuhao, et autres
Publié: (2024)
par: Du, Yuhao, et autres
Publié: (2024)
Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models
par: Yu, Miao, et autres
Publié: (2025)
par: Yu, Miao, et autres
Publié: (2025)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
par: Tong, Terry, et autres
Publié: (2024)
par: Tong, Terry, et autres
Publié: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
par: Ying, Zonghao, et autres
Publié: (2025)
par: Ying, Zonghao, et autres
Publié: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
par: Zhao, Shuai, et autres
Publié: (2024)
par: Zhao, Shuai, et autres
Publié: (2024)
Documents similaires
-
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space
par: Wu, Zongru, et autres
Publié: (2024) -
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
par: Cheng, Pengzhou, et autres
Publié: (2024) -
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models
par: Cheng, Pengzhou, et autres
Publié: (2024) -
MKF-ADS: Multi-Knowledge Fusion Based Self-supervised Anomaly Detection System for Control Area Network
par: Cheng, Pengzhou, et autres
Publié: (2024) -
Transferring Backdoors between Large Language Models by Knowledge Distillation
par: Cheng, Pengzhou, et autres
Publié: (2024)