CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuetai, Xu, Zhangchen, Jiang, Fengqing, Niu, Luyao, Sahabandu, Dinuka, Ramasubramanian, Bhaskar, Poovendran, Radha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
by: Sahabandu, Dinuka, et al.
Published: (2024)
by: Sahabandu, Dinuka, et al.
Published: (2024)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024)
by: Xiang, Zhen, et al.
Published: (2024)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Who is Responsible? Explaining Safety Violations in Multi-Agent Cyber-Physical Systems
by: Niu, Luyao, et al.
Published: (2024)
by: Niu, Luyao, et al.
Published: (2024)
CANTXSec: A Deterministic Intrusion Detection and Prevention System for CAN Bus Monitoring ECU Activations
by: Donadel, Denis, et al.
Published: (2025)
by: Donadel, Denis, et al.
Published: (2025)
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
Double-Dip: Thwarting Label-Only Membership Inference Attacks with Transfer Learning and Randomization
by: Rajabi, Arezoo, et al.
Published: (2024)
by: Rajabi, Arezoo, et al.
Published: (2024)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
by: Feng, Yichen, et al.
Published: (2025)
by: Feng, Yichen, et al.
Published: (2025)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
by: Jiang, Fengqing, et al.
Published: (2024)
by: Jiang, Fengqing, et al.
Published: (2024)
Temporal Sampling for Forgotten Reasoning in LLMs
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Small Models Struggle to Learn from Strong Reasoners
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
by: Zhang, Jiale, et al.
Published: (2025)
by: Zhang, Jiale, et al.
Published: (2025)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Clean-image Backdoor Attacks
by: Rong, Dazhong, et al.
Published: (2024)
by: Rong, Dazhong, et al.
Published: (2024)
A Method for Fast Autonomy Transfer in Reinforcement Learning
by: Sahabandu, Dinuka, et al.
Published: (2024)
by: Sahabandu, Dinuka, et al.
Published: (2024)
Clean-Label Physical Backdoor Attacks with Data Distillation
by: Dao, Thinh, et al.
Published: (2024)
by: Dao, Thinh, et al.
Published: (2024)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
Generalization Bound and New Algorithm for Clean-Label Backdoor Attack
by: Yu, Lijia, et al.
Published: (2024)
by: Yu, Lijia, et al.
Published: (2024)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models
by: Fu, Hang, et al.
Published: (2026)
by: Fu, Hang, et al.
Published: (2026)
UltraClean: A Simple Framework to Train Robust Neural Networks against Backdoor Attacks
by: Zhao, Bingyin, et al.
Published: (2023)
by: Zhao, Bingyin, et al.
Published: (2023)
Selection-Based Vulnerabilities: Clean-Label Backdoor Attacks in Active Learning
by: Zhi, Yuhan, et al.
Published: (2025)
by: Zhi, Yuhan, et al.
Published: (2025)
FFCBA: Feature-based Full-target Clean-label Backdoor Attacks
by: Yin, Yangxu, et al.
Published: (2025)
by: Yin, Yangxu, et al.
Published: (2025)
Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models
by: Hao, Jiang, et al.
Published: (2024)
by: Hao, Jiang, et al.
Published: (2024)
One-to-Multiple Clean-Label Image Camouflage (OmClic) based Backdoor Attack on Deep Learning
by: Wang, Guohong, et al.
Published: (2023)
by: Wang, Guohong, et al.
Published: (2023)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
by: Jiang, Fengqing, et al.
Published: (2025)
by: Jiang, Fengqing, et al.
Published: (2025)
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
by: Hossen, Md Imran, et al.
Published: (2024)
by: Hossen, Md Imran, et al.
Published: (2024)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
by: Ding, Yidong, et al.
Published: (2025)
by: Ding, Yidong, et al.
Published: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Mitigating Backdoor Triggered and Targeted Data Poisoning Attacks in Voice Authentication Systems
by: Mohammadi, Alireza, et al.
Published: (2025)
by: Mohammadi, Alireza, et al.
Published: (2025)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
by: Xu, Zhangchen, et al.
Published: (2024)
by: Xu, Zhangchen, et al.
Published: (2024)
Mitigating Backdoor Attack by Injecting Proactive Defensive Backdoor
by: Wei, Shaokui, et al.
Published: (2024)
by: Wei, Shaokui, et al.
Published: (2024)
BadToken: Token-level Backdoor Attacks to Multi-modal Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
The WidthWall: A Strict Expressivity Hierarchy for Hypergraph Neural Networks
by: Jiang, Fengqing, et al.
Published: (2026)
by: Jiang, Fengqing, et al.
Published: (2026)
Similar Items
-
Game of Trojans: Adaptive Adversaries Against Output-based Trojaned-Model Detectors
by: Sahabandu, Dinuka, et al.
Published: (2024) -
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
by: Xiang, Zhen, et al.
Published: (2024) -
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024) -
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
by: Jiang, Fengqing, et al.
Published: (2025) -
Brave: Byzantine-Resilient and Privacy-Preserving Peer-to-Peer Federated Learning
by: Xu, Zhangchen, et al.
Published: (2024)