Task-Agnostic Detector for Insertion-Based Backdoor Attacks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lyu, Weimin, Lin, Xiao, Zheng, Songzhu, Pang, Lu, Ling, Haibin, Jha, Susmit, Chen, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Long-Tailed Backdoor Attack Using Dynamic Data Augmentation Operations
von: Pang, Lu, et al.
Veröffentlicht: (2024)
von: Pang, Lu, et al.
Veröffentlicht: (2024)
Test-Time Backdoor Attacks on Multimodal Large Language Models
von: Lu, Dong, et al.
Veröffentlicht: (2024)
von: Lu, Dong, et al.
Veröffentlicht: (2024)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
von: Wen, Rui, et al.
Veröffentlicht: (2026)
von: Wen, Rui, et al.
Veröffentlicht: (2026)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024)
CAPAA: Classifier-Agnostic Projector-Based Adversarial Attack
von: Li, Zhan, et al.
Veröffentlicht: (2025)
von: Li, Zhan, et al.
Veröffentlicht: (2025)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
von: Cheng, Pengzhou, et al.
Veröffentlicht: (2024)
Denial-of-Service Poisoning Attacks against Large Language Models
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
von: Gao, Kuofeng, et al.
Veröffentlicht: (2024)
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
von: Peng, Yuefeng, et al.
Veröffentlicht: (2024)
BadApex: Backdoor Attack Based on Adaptive Optimization Mechanism of Black-box Large Language Models
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2025)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Large Language Models are Good Attackers: Efficient and Stealthy Textual Backdoor Attacks
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
von: Li, Ziqiang, et al.
Veröffentlicht: (2024)
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
von: He, Xuanli, et al.
Veröffentlicht: (2024)
von: He, Xuanli, et al.
Veröffentlicht: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
von: He, Xuanli, et al.
Veröffentlicht: (2024)
von: He, Xuanli, et al.
Veröffentlicht: (2024)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
von: Jiang, Peihai, et al.
Veröffentlicht: (2025)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
von: Wei, Jiali, et al.
Veröffentlicht: (2026)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
von: Zhang, Xiaozhe, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaozhe, et al.
Veröffentlicht: (2026)
Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
von: Wang, Tianchun, et al.
Veröffentlicht: (2024)
BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xue, Jiaqi, et al.
Veröffentlicht: (2024)
Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
von: Li, Haoran, et al.
Veröffentlicht: (2024)
von: Li, Haoran, et al.
Veröffentlicht: (2024)
Claim-Guided Textual Backdoor Attack for Practical Applications
von: Song, Minkyoo, et al.
Veröffentlicht: (2024)
von: Song, Minkyoo, et al.
Veröffentlicht: (2024)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Concept-Guided Backdoor Attack on Vision Language Models
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
von: Shen, Haoyu, et al.
Veröffentlicht: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
von: Chaudhari, Harsh, et al.
Veröffentlicht: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
von: Du, Wei, et al.
Veröffentlicht: (2023)
von: Du, Wei, et al.
Veröffentlicht: (2023)
GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
MBTSAD: Mitigating Backdoors in Language Models Based on Token Splitting and Attention Distillation
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
von: Ding, Yidong, et al.
Veröffentlicht: (2025)
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
von: Mangaokar, Neal, et al.
Veröffentlicht: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
von: Zhao, Shuai, et al.
Veröffentlicht: (2024)
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
von: Dong, Yiting, et al.
Veröffentlicht: (2024)
von: Dong, Yiting, et al.
Veröffentlicht: (2024)
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
von: Yang, Wenkai, et al.
Veröffentlicht: (2024)
BadActs: A Universal Backdoor Defense in the Activation Space
von: Yi, Biao, et al.
Veröffentlicht: (2024)
von: Yi, Biao, et al.
Veröffentlicht: (2024)
Privacy Preserving In-Context-Learning Framework for Large Language Models
von: Bhusal, Bishnu, et al.
Veröffentlicht: (2025)
von: Bhusal, Bishnu, et al.
Veröffentlicht: (2025)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
von: Li, Yuanfan, et al.
Veröffentlicht: (2026)
von: Li, Yuanfan, et al.
Veröffentlicht: (2026)
IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
von: Li, Junxian, et al.
Veröffentlicht: (2025)
von: Li, Junxian, et al.
Veröffentlicht: (2025)
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Organization Structures
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
von: Li, Bangxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Long-Tailed Backdoor Attack Using Dynamic Data Augmentation Operations
von: Pang, Lu, et al.
Veröffentlicht: (2024) -
Test-Time Backdoor Attacks on Multimodal Large Language Models
von: Lu, Dong, et al.
Veröffentlicht: (2024) -
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
von: Wen, Rui, et al.
Veröffentlicht: (2026) -
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
von: Wang, Jiongxiao, et al.
Veröffentlicht: (2024) -
CAPAA: Classifier-Agnostic Projector-Based Adversarial Attack
von: Li, Zhan, et al.
Veröffentlicht: (2025)