One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Linbao, Liu, Yannan, He, Daojing, Li, Yu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
par: Chen, Yunhao, et autres
Publié: (2025)
par: Chen, Yunhao, et autres
Publié: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
par: Zhang, Chiyu, et autres
Publié: (2025)
par: Zhang, Chiyu, et autres
Publié: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
par: Li, Xiang, et autres
Publié: (2025)
par: Li, Xiang, et autres
Publié: (2025)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
par: Ostermann, Simon, et autres
Publié: (2024)
par: Ostermann, Simon, et autres
Publié: (2024)
Proactive defense against LLM Jailbreak
par: Zhao, Weiliang, et autres
Publié: (2025)
par: Zhao, Weiliang, et autres
Publié: (2025)
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
par: Yu, Zhiyuan, et autres
Publié: (2024)
par: Yu, Zhiyuan, et autres
Publié: (2024)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
par: Luo, Xuan, et autres
Publié: (2025)
par: Luo, Xuan, et autres
Publié: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
par: Shen, Guobin, et autres
Publié: (2025)
par: Shen, Guobin, et autres
Publié: (2025)
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
par: Xie, Yueqi, et autres
Publié: (2024)
par: Xie, Yueqi, et autres
Publié: (2024)
Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs
par: Liu, Yize, et autres
Publié: (2025)
par: Liu, Yize, et autres
Publié: (2025)
Imperceptible Jailbreaking against Large Language Models
par: Gao, Kuofeng, et autres
Publié: (2025)
par: Gao, Kuofeng, et autres
Publié: (2025)
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
par: Pathade, Chetan
Publié: (2025)
par: Pathade, Chetan
Publié: (2025)
RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction
par: Jiang, Tanqiu, et autres
Publié: (2024)
par: Jiang, Tanqiu, et autres
Publié: (2024)
Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
par: Hu, Hanjiang, et autres
Publié: (2025)
par: Hu, Hanjiang, et autres
Publié: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
par: Yeo, Wei Jie, et autres
Publié: (2025)
par: Yeo, Wei Jie, et autres
Publié: (2025)
PRSA: Prompt Stealing Attacks against Real-World Prompt Services
par: Yang, Yong, et autres
Publié: (2024)
par: Yang, Yong, et autres
Publié: (2024)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
par: Kirch, Nathalie, et autres
Publié: (2024)
par: Kirch, Nathalie, et autres
Publié: (2024)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
par: Li, Yucheng, et autres
Publié: (2025)
par: Li, Yucheng, et autres
Publié: (2025)
Are All Prompt Components Value-Neutral? Understanding the Heterogeneous Adversarial Robustness of Dissected Prompt in Large Language Models
par: Zheng, Yujia, et autres
Publié: (2025)
par: Zheng, Yujia, et autres
Publié: (2025)
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance
par: Huang, Caishuang, et autres
Publié: (2024)
par: Huang, Caishuang, et autres
Publié: (2024)
Defending against Jailbreak through Early Exit Generation of Large Language Models
par: Zhao, Chongwen, et autres
Publié: (2024)
par: Zhao, Chongwen, et autres
Publié: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
par: Liu, Songyang, et autres
Publié: (2026)
par: Liu, Songyang, et autres
Publié: (2026)
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
par: Luo, Weidi, et autres
Publié: (2024)
par: Luo, Weidi, et autres
Publié: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
par: Ji, Wence, et autres
Publié: (2025)
par: Ji, Wence, et autres
Publié: (2025)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
par: Wu, Yu-Hang, et autres
Publié: (2025)
par: Wu, Yu-Hang, et autres
Publié: (2025)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
par: He, Yu, et autres
Publié: (2025)
par: He, Yu, et autres
Publié: (2025)
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models
par: Wang, Yiming, et autres
Publié: (2024)
par: Wang, Yiming, et autres
Publié: (2024)
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
par: Xiong, Chen, et autres
Publié: (2024)
par: Xiong, Chen, et autres
Publié: (2024)
Fingerprinting LLMs via Prompt Injection
par: Hu, Yuepeng, et autres
Publié: (2025)
par: Hu, Yuepeng, et autres
Publié: (2025)
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
par: Xu, Zhao, et autres
Publié: (2024)
par: Xu, Zhao, et autres
Publié: (2024)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
par: Liang, Buyun, et autres
Publié: (2025)
par: Liang, Buyun, et autres
Publié: (2025)
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
par: Tshimula, Jean Marie, et autres
Publié: (2024)
par: Tshimula, Jean Marie, et autres
Publié: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
par: Li, Qizhang, et autres
Publié: (2024)
par: Li, Qizhang, et autres
Publié: (2024)
Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs
par: Pu, Rui, et autres
Publié: (2024)
par: Pu, Rui, et autres
Publié: (2024)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
par: Liu, Fan, et autres
Publié: (2024)
par: Liu, Fan, et autres
Publié: (2024)
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
par: Dong, Yiting, et autres
Publié: (2024)
par: Dong, Yiting, et autres
Publié: (2024)
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters
par: Yang, Yan, et autres
Publié: (2024)
par: Yang, Yan, et autres
Publié: (2024)
Jailbreaking with Universal Multi-Prompts
par: Hsu, Yu-Ling, et autres
Publié: (2025)
par: Hsu, Yu-Ling, et autres
Publié: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
par: Li, Xirui, et autres
Publié: (2024)
par: Li, Xirui, et autres
Publié: (2024)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
par: Xu, Ning, et autres
Publié: (2025)
par: Xu, Ning, et autres
Publié: (2025)
Documents similaires
-
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
par: Chen, Yunhao, et autres
Publié: (2025) -
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
par: Zhang, Chiyu, et autres
Publié: (2025) -
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
par: Li, Xiang, et autres
Publié: (2025) -
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
par: Ostermann, Simon, et autres
Publié: (2024) -
Proactive defense against LLM Jailbreak
par: Zhao, Weiliang, et autres
Publié: (2025)