Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Shuai, Wen, Jinming, Tuan, Luu Anh, Zhao, Junbo, Fu, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
Aspect-Based Summarization with Self-Aspect Retrieval Enhanced Generation
by: Feng, Yichao, et al.
Published: (2025)
by: Feng, Yichao, et al.
Published: (2025)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
by: Srey, Ponhvoan, et al.
Published: (2026)
by: Srey, Ponhvoan, et al.
Published: (2026)
Learning Uncertainty from Sequential Internal Dispersion in Large Language Models
by: Srey, Ponhvoan, et al.
Published: (2026)
by: Srey, Ponhvoan, et al.
Published: (2026)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models
by: Liu, Ziyu, et al.
Published: (2026)
by: Liu, Ziyu, et al.
Published: (2026)
Unsupervised Hallucination Detection by Inspecting Reasoning Processes
by: Srey, Ponhvoan, et al.
Published: (2025)
by: Srey, Ponhvoan, et al.
Published: (2025)
A Survey on Neural Topic Models: Methods, Applications, and Challenges
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
Don't Forget Your Reward Values: Language Model Alignment via Value-based Calibration
by: Mao, Xin, et al.
Published: (2024)
by: Mao, Xin, et al.
Published: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023)
by: Du, Wei, et al.
Published: (2023)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
by: Pham, Quang Hieu, et al.
Published: (2024)
by: Pham, Quang Hieu, et al.
Published: (2024)
Towards the TopMost: A Topic Modeling System Toolkit
by: Wu, Xiaobao, et al.
Published: (2023)
by: Wu, Xiaobao, et al.
Published: (2023)
Modeling Dynamic Topics in Chain-Free Fashion by Evolution-Tracking Contrastive Learning and Unassociated Word Exclusion
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models
by: Kim, San, et al.
Published: (2026)
by: Kim, San, et al.
Published: (2026)
Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
by: Ding, Bosheng, et al.
Published: (2024)
by: Ding, Bosheng, et al.
Published: (2024)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
AKEW: Assessing Knowledge Editing in the Wild
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model
by: Wu, Xiaobao, et al.
Published: (2024)
by: Wu, Xiaobao, et al.
Published: (2024)
An Examination on the Effectiveness of Divide-and-Conquer Prompting in Large Language Models
by: Zhang, Yizhou, et al.
Published: (2024)
by: Zhang, Yizhou, et al.
Published: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
by: Hao, Yunzhuo, et al.
Published: (2024)
by: Hao, Yunzhuo, et al.
Published: (2024)
Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
by: Lin, Tzu-Ling, et al.
Published: (2025)
by: Lin, Tzu-Ling, et al.
Published: (2025)
BadCLM: Backdoor Attack in Clinical Language Models for Electronic Health Records
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
MEGen: Generative Backdoor into Large Language Models via Model Editing
by: Qiu, Jiyang, et al.
Published: (2024)
by: Qiu, Jiyang, et al.
Published: (2024)
Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in Large Language Models
by: Hu, Zhiyuan, et al.
Published: (2024)
by: Hu, Zhiyuan, et al.
Published: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
by: Van Long, Phuoc Pham, et al.
Published: (2023)
by: Van Long, Phuoc Pham, et al.
Published: (2023)
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks
by: Mu, Lin, et al.
Published: (2025)
by: Mu, Lin, et al.
Published: (2025)
Metacognitive Prompting Improves Understanding in Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Similar Items
-
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024) -
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024) -
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024) -
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024) -
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)