Instruction Backdoor Attacks Against Customized LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Rui, Li, Hongwei, Wen, Rui, Jiang, Wenbo, Zhang, Yuan, Backes, Michael, Shen, Yun, Zhang, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
von: Wen, Rui, et al.
Veröffentlicht: (2025)
von: Wen, Rui, et al.
Veröffentlicht: (2025)
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
von: Shen, Xinyue, et al.
Veröffentlicht: (2024)
Prompt Stealing Attacks Against Text-to-Image Generation Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
BadMerging: Backdoor Attacks Against Model Merging
von: Zhang, Jinghuai, et al.
Veröffentlicht: (2024)
von: Zhang, Jinghuai, et al.
Veröffentlicht: (2024)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
von: Yang, Ziqing, et al.
Veröffentlicht: (2026)
von: Yang, Ziqing, et al.
Veröffentlicht: (2026)
Excessive Reasoning Attack on Reasoning LLMs
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
von: Si, Wai Man, et al.
Veröffentlicht: (2025)
BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
von: Wang, Zihan, et al.
Veröffentlicht: (2026)
On the Out-of-Distribution Backdoor Attack for Federated Learning
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
Combinational Backdoor Attack against Customized Text-to-Image Models
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
von: Jiang, Wenbo, et al.
Veröffentlicht: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
von: Chu, Junjie, et al.
Veröffentlicht: (2024)
Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
von: Zhu, Rui, et al.
Veröffentlicht: (2023)
von: Zhu, Rui, et al.
Veröffentlicht: (2023)
When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTs
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
von: Shen, Xinyue, et al.
Veröffentlicht: (2025)
Membership Inference Attacks Against In-Context Learning
von: Wen, Rui, et al.
Veröffentlicht: (2024)
von: Wen, Rui, et al.
Veröffentlicht: (2024)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Vera Verto: Multimodal Hijacking Attack
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2023)
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
von: Guo, Zhen, et al.
Veröffentlicht: (2025)
von: Guo, Zhen, et al.
Veröffentlicht: (2025)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
Link Stealing Attacks Against Inductive Graph Neural Networks
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
von: Wu, Yixin, et al.
Veröffentlicht: (2024)
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
von: Guo, Ji, et al.
Veröffentlicht: (2026)
von: Guo, Ji, et al.
Veröffentlicht: (2026)
Transferable Availability Poisoning Attacks
von: Liu, Yiyong, et al.
Veröffentlicht: (2023)
von: Liu, Yiyong, et al.
Veröffentlicht: (2023)
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
von: Wang, Yifei, et al.
Veröffentlicht: (2026)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
von: Zhang, Rui, et al.
Veröffentlicht: (2026)
Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
von: Zhang, Minxing, et al.
Veröffentlicht: (2024)
Heterogeneous Graph Backdoor Attack
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
The Challenge of Identifying the Origin of Black-Box Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
von: Yang, Ziqing, et al.
Veröffentlicht: (2025)
UIBDiffusion: Universal Imperceptible Backdoor Attack for Diffusion Models
von: Han, Yuning, et al.
Veröffentlicht: (2024)
von: Han, Yuning, et al.
Veröffentlicht: (2024)
Krait: A Backdoor Attack Against Graph Prompt Tuning
von: Song, Ying, et al.
Veröffentlicht: (2024)
von: Song, Ying, et al.
Veröffentlicht: (2024)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
von: Jiang, Yukun, et al.
Veröffentlicht: (2026)
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
Compromising Embodied Agents with Contextual Backdoor Attacks
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
von: Liu, Aishan, et al.
Veröffentlicht: (2024)
Detecting Backdoor Attacks in Federated Learning via Direction Alignment Inspection
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
von: Xu, Jiahao, et al.
Veröffentlicht: (2025)
Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
von: Wu, Yixin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023) -
SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark
von: Wen, Rui, et al.
Veröffentlicht: (2025) -
Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?
von: Wen, Rui, et al.
Veröffentlicht: (2024) -
Voice Jailbreak Attacks Against GPT-4o
von: Shen, Xinyue, et al.
Veröffentlicht: (2024) -
Prompt Stealing Attacks Against Text-to-Image Generation Models
von: Shen, Xinyue, et al.
Veröffentlicht: (2023)