FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Bocheng, Guo, Hanqing, Yan, Qiben |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Boundaries: A Comprehensive Survey of Transferable Attacks on AI Systems
por: Wang, Guangjing, et al.
Publicado: (2023)
por: Wang, Guangjing, et al.
Publicado: (2023)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
por: Chen, Bocheng, et al.
Publicado: (2024)
por: Chen, Bocheng, et al.
Publicado: (2024)
ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
por: Wang, Yuanda, et al.
Publicado: (2025)
por: Wang, Yuanda, et al.
Publicado: (2025)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
por: Akbar-Tajari, Mohammad, et al.
Publicado: (2025)
por: Akbar-Tajari, Mohammad, et al.
Publicado: (2025)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
por: Gohil, Vasudev
Publicado: (2025)
por: Gohil, Vasudev
Publicado: (2025)
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
por: Mehrotra, Anay, et al.
Publicado: (2023)
por: Mehrotra, Anay, et al.
Publicado: (2023)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
por: Zhang, Chiyu, et al.
Publicado: (2025)
por: Zhang, Chiyu, et al.
Publicado: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
por: Li, Yucheng, et al.
Publicado: (2025)
por: Li, Yucheng, et al.
Publicado: (2025)
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
por: Hu, Xiaomeng, et al.
Publicado: (2025)
por: Hu, Xiaomeng, et al.
Publicado: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)
por: Liu, Fan, et al.
Publicado: (2024)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
por: Kim, Heegyu, et al.
Publicado: (2024)
por: Kim, Heegyu, et al.
Publicado: (2024)
Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks
por: Liu, Xiaoqun, et al.
Publicado: (2024)
por: Liu, Xiaoqun, et al.
Publicado: (2024)
ShieldLearner: A New Paradigm for Jailbreak Attack Defense in LLMs
por: Ni, Ziyi, et al.
Publicado: (2025)
por: Ni, Ziyi, et al.
Publicado: (2025)
Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
por: Zhong, Xingwei, et al.
Publicado: (2025)
por: Zhong, Xingwei, et al.
Publicado: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025)
por: Chen, Yunhao, et al.
Publicado: (2025)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
por: Zhang, Xiaozhe, et al.
Publicado: (2026)
por: Zhang, Xiaozhe, et al.
Publicado: (2026)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
por: Shen, Guobin, et al.
Publicado: (2025)
por: Shen, Guobin, et al.
Publicado: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
por: Chu, Junjie, et al.
Publicado: (2024)
por: Chu, Junjie, et al.
Publicado: (2024)
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
por: Yan, Dong, et al.
Publicado: (2026)
por: Yan, Dong, et al.
Publicado: (2026)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
por: Yi, Sibo, et al.
Publicado: (2024)
por: Yi, Sibo, et al.
Publicado: (2024)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
por: Brown, Hannah, et al.
Publicado: (2024)
por: Brown, Hannah, et al.
Publicado: (2024)
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
por: Feng, Yingchaojie, et al.
Publicado: (2024)
por: Feng, Yingchaojie, et al.
Publicado: (2024)
Black-Box Guardrail Reverse-engineering Attack
por: Yao, Hongwei, et al.
Publicado: (2025)
por: Yao, Hongwei, et al.
Publicado: (2025)
LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
por: Li, Nathaniel, et al.
Publicado: (2024)
por: Li, Nathaniel, et al.
Publicado: (2024)
Proactive defense against LLM Jailbreak
por: Zhao, Weiliang, et al.
Publicado: (2025)
por: Zhao, Weiliang, et al.
Publicado: (2025)
Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework
por: Wang, Zhuoshang, et al.
Publicado: (2026)
por: Wang, Zhuoshang, et al.
Publicado: (2026)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
por: Afane, Mohamed, et al.
Publicado: (2025)
por: Afane, Mohamed, et al.
Publicado: (2025)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
por: Yu, Miao, et al.
Publicado: (2024)
por: Yu, Miao, et al.
Publicado: (2024)
A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy
por: Correia, Pedro H. Barcha, et al.
Publicado: (2026)
por: Correia, Pedro H. Barcha, et al.
Publicado: (2026)
Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning
por: An, Li, et al.
Publicado: (2025)
por: An, Li, et al.
Publicado: (2025)
PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
por: Jawad, Huseein, et al.
Publicado: (2025)
por: Jawad, Huseein, et al.
Publicado: (2025)
DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation
por: Jiang, Bo
Publicado: (2026)
por: Jiang, Bo
Publicado: (2026)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
por: Wu, Yu-Hang, et al.
Publicado: (2025)
por: Wu, Yu-Hang, et al.
Publicado: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
por: Li, Xirui, et al.
Publicado: (2024)
por: Li, Xirui, et al.
Publicado: (2024)
PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs
por: Wang, Xinkai, et al.
Publicado: (2025)
por: Wang, Xinkai, et al.
Publicado: (2025)
"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks
por: Wang, Libo
Publicado: (2024)
por: Wang, Libo
Publicado: (2024)
AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs
por: Lv, Lijia, et al.
Publicado: (2024)
por: Lv, Lijia, et al.
Publicado: (2024)
Ejemplares similares
-
Beyond Boundaries: A Comprehensive Survey of Transferable Attacks on AI Systems
por: Wang, Guangjing, et al.
Publicado: (2023) -
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
por: Chen, Bocheng, et al.
Publicado: (2024) -
ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
por: Wang, Yuanda, et al.
Publicado: (2025) -
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024) -
Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
por: Akbar-Tajari, Mohammad, et al.
Publicado: (2025)