ASPIRER: Bypassing System Prompts With Permutation-based Backdoors in LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Lu, Cheng, Siyuan, Chen, Xuan, Zhang, Kaiyuan, Shen, Guangyu, Zhang, Zhuo, Zhang, Xiangyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
por: Shen, Guangyu, et al.
Publicado: (2025)
por: Shen, Guangyu, et al.
Publicado: (2025)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
por: Shen, Guangyu, et al.
Publicado: (2024)
por: Shen, Guangyu, et al.
Publicado: (2024)
UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening
por: Cheng, Siyuan, et al.
Publicado: (2024)
por: Cheng, Siyuan, et al.
Publicado: (2024)
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
por: Yan, Lu, et al.
Publicado: (2025)
por: Yan, Lu, et al.
Publicado: (2025)
Opening A Pandora's Box: Things You Should Know in the Era of Custom GPTs
por: Tao, Guanhong, et al.
Publicado: (2023)
por: Tao, Guanhong, et al.
Publicado: (2023)
LOTUS: Evasive and Resilient Backdoor Attacks through Sub-Partitioning
por: Cheng, Siyuan, et al.
Publicado: (2024)
por: Cheng, Siyuan, et al.
Publicado: (2024)
CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
por: An, Shengwei, et al.
Publicado: (2023)
por: An, Shengwei, et al.
Publicado: (2023)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
por: Lu, Yiyang, et al.
Publicado: (2026)
por: Lu, Yiyang, et al.
Publicado: (2026)
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMs
por: Chen, Xuan, et al.
Publicado: (2024)
por: Chen, Xuan, et al.
Publicado: (2024)
Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving
por: Chen, Xuan, et al.
Publicado: (2025)
por: Chen, Xuan, et al.
Publicado: (2025)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
por: Chen, Xuan, et al.
Publicado: (2026)
por: Chen, Xuan, et al.
Publicado: (2026)
Instruction Backdoor Attacks Against Customized LLMs
por: Zhang, Rui, et al.
Publicado: (2024)
por: Zhang, Rui, et al.
Publicado: (2024)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
por: Xu, Xiangzhe, et al.
Publicado: (2025)
por: Xu, Xiangzhe, et al.
Publicado: (2025)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
por: Yang, Guang, et al.
Publicado: (2024)
por: Yang, Guang, et al.
Publicado: (2024)
The Invitation Trap: Proactive Availability Backdoor in LLMs via Conversational Induction
por: Wang, He, et al.
Publicado: (2026)
por: Wang, He, et al.
Publicado: (2026)
PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification
por: Gong, Guangyu, et al.
Publicado: (2026)
por: Gong, Guangyu, et al.
Publicado: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
por: Guo, Weiyang, et al.
Publicado: (2026)
por: Guo, Weiyang, et al.
Publicado: (2026)
Backdoor Contrastive Learning via Bi-level Trigger Optimization
por: Sun, Weiyu, et al.
Publicado: (2024)
por: Sun, Weiyu, et al.
Publicado: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
A Practical Trigger-Free Backdoor Attack on Neural Networks
por: Wang, Jiahao, et al.
Publicado: (2024)
por: Wang, Jiahao, et al.
Publicado: (2024)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
por: Xu, Xiangzhe, et al.
Publicado: (2024)
por: Xu, Xiangzhe, et al.
Publicado: (2024)
SSD: A State-based Stealthy Backdoor Attack For Navigation System in UAV Route Planning
por: Wang, Zhaoxuan, et al.
Publicado: (2025)
por: Wang, Zhaoxuan, et al.
Publicado: (2025)
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search
por: Chen, Xuan, et al.
Publicado: (2024)
por: Chen, Xuan, et al.
Publicado: (2024)
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
por: Wang, Xuan, et al.
Publicado: (2025)
por: Wang, Xuan, et al.
Publicado: (2025)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
por: Chen, Chen, et al.
Publicado: (2026)
por: Chen, Chen, et al.
Publicado: (2026)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
por: Chen, Yulin, et al.
Publicado: (2025)
por: Chen, Yulin, et al.
Publicado: (2025)
The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
por: Wei, Rongzhe, et al.
Publicado: (2025)
por: Wei, Rongzhe, et al.
Publicado: (2025)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
por: Chen, Zhuowei, et al.
Publicado: (2025)
por: Chen, Zhuowei, et al.
Publicado: (2025)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
por: Zhang, Rui, et al.
Publicado: (2025)
por: Zhang, Rui, et al.
Publicado: (2025)
Bypassing Prompt Guards in Production with Controlled-Release Prompting
por: Fairoze, Jaiden, et al.
Publicado: (2025)
por: Fairoze, Jaiden, et al.
Publicado: (2025)
Imperceptible Sample-Specific Backdoor to DNN with Denoising Autoencoder
por: Wang, Xiangqi, et al.
Publicado: (2023)
por: Wang, Xiangqi, et al.
Publicado: (2023)
STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling
por: Wang, Kun, et al.
Publicado: (2026)
por: Wang, Kun, et al.
Publicado: (2026)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
por: Cui, Jing, et al.
Publicado: (2025)
por: Cui, Jing, et al.
Publicado: (2025)
LLM Agents Should Employ Security Principles
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
Activation Gradient based Poisoned Sample Detection Against Backdoor Attacks
por: Yuan, Danni, et al.
Publicado: (2023)
por: Yuan, Danni, et al.
Publicado: (2023)
Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning
por: Zhang, Bokang, et al.
Publicado: (2025)
por: Zhang, Bokang, et al.
Publicado: (2025)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
por: Zhang, Xiaoyu, et al.
Publicado: (2023)
Ejemplares similares
-
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
por: Shen, Guangyu, et al.
Publicado: (2025) -
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
por: Shen, Guangyu, et al.
Publicado: (2024) -
UNIT: Backdoor Mitigation via Automated Neural Distribution Tightening
por: Cheng, Siyuan, et al.
Publicado: (2024) -
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
por: Yan, Lu, et al.
Publicado: (2025) -
Opening A Pandora's Box: Things You Should Know in the Era of Custom GPTs
por: Tao, Guanhong, et al.
Publicado: (2023)