SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jindong, Liu, Ying, Fu, Yali, Zhu, Jinjing, Wang, Leyao, Yang, Menglin, Ying, Rex |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM Jailbreak Detection for (Almost) Free!
por: Chen, Guorui, et al.
Publicado: (2025)
por: Chen, Guorui, et al.
Publicado: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025)
por: Chen, Yunhao, et al.
Publicado: (2025)
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
por: Shen, Guobin, et al.
Publicado: (2025)
por: Shen, Guobin, et al.
Publicado: (2025)
Proactive defense against LLM Jailbreak
por: Zhao, Weiliang, et al.
Publicado: (2025)
por: Zhao, Weiliang, et al.
Publicado: (2025)
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
por: Zhang, Junke, et al.
Publicado: (2026)
por: Zhang, Junke, et al.
Publicado: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
por: Huang, Ruixuan, et al.
Publicado: (2025)
por: Huang, Ruixuan, et al.
Publicado: (2025)
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
por: Zhou, Yukai, et al.
Publicado: (2025)
por: Zhou, Yukai, et al.
Publicado: (2025)
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
por: Zhao, Weixiang, et al.
Publicado: (2025)
por: Zhao, Weixiang, et al.
Publicado: (2025)
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
por: Zhang, Chiyu, et al.
Publicado: (2026)
por: Zhang, Chiyu, et al.
Publicado: (2026)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
por: Jia, Xiaojun, et al.
Publicado: (2024)
por: Jia, Xiaojun, et al.
Publicado: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
por: Liu, Songyang, et al.
Publicado: (2026)
por: Liu, Songyang, et al.
Publicado: (2026)
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
por: Dou, Yipu, et al.
Publicado: (2026)
por: Dou, Yipu, et al.
Publicado: (2026)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
por: Li, Yucheng, et al.
Publicado: (2025)
por: Li, Yucheng, et al.
Publicado: (2025)
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
por: Wu, Yu-Hang, et al.
Publicado: (2025)
por: Wu, Yu-Hang, et al.
Publicado: (2025)
TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
por: Zeng, Churui, et al.
Publicado: (2026)
por: Zeng, Churui, et al.
Publicado: (2026)
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
por: Wang, Zhaoqi, et al.
Publicado: (2025)
por: Wang, Zhaoqi, et al.
Publicado: (2025)
FlexLLM: Exploring LLM Customization for Moving Target Defense on Black-Box LLMs Against Jailbreak Attacks
por: Chen, Bocheng, et al.
Publicado: (2024)
por: Chen, Bocheng, et al.
Publicado: (2024)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
por: Li, Linbao, et al.
Publicado: (2025)
por: Li, Linbao, et al.
Publicado: (2025)
Geneshift: Impact of different scenario shift on Jailbreaking LLM
por: Wu, Tianyi, et al.
Publicado: (2025)
por: Wu, Tianyi, et al.
Publicado: (2025)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
por: Li, Xirui, et al.
Publicado: (2024)
por: Li, Xirui, et al.
Publicado: (2024)
The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis
por: Wang, Peiran, et al.
Publicado: (2026)
por: Wang, Peiran, et al.
Publicado: (2026)
SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
por: Xin, Yuan, et al.
Publicado: (2026)
por: Xin, Yuan, et al.
Publicado: (2026)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
por: Assogba, Yannick, et al.
Publicado: (2026)
por: Assogba, Yannick, et al.
Publicado: (2026)
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
por: Yu, Miao, et al.
Publicado: (2024)
por: Yu, Miao, et al.
Publicado: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
por: You, Wenhao, et al.
Publicado: (2025)
por: You, Wenhao, et al.
Publicado: (2025)
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
por: Wang, Jiongxiao, et al.
Publicado: (2024)
por: Wang, Jiongxiao, et al.
Publicado: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
por: Zhu, Sicheng, et al.
Publicado: (2024)
por: Zhu, Sicheng, et al.
Publicado: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
por: Das, Saswat, et al.
Publicado: (2025)
por: Das, Saswat, et al.
Publicado: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
por: Yeo, Wei Jie, et al.
Publicado: (2025)
por: Yeo, Wei Jie, et al.
Publicado: (2025)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
por: Zhang, Chiyu, et al.
Publicado: (2025)
por: Zhang, Chiyu, et al.
Publicado: (2025)
Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
por: Tu, Shangqing, et al.
Publicado: (2024)
por: Tu, Shangqing, et al.
Publicado: (2024)
Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2026)
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
por: Yadav, Advait, et al.
Publicado: (2025)
por: Yadav, Advait, et al.
Publicado: (2025)
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?
por: Chen, Shuo, et al.
Publicado: (2024)
por: Chen, Shuo, et al.
Publicado: (2024)
Ejemplares similares
-
LLM Jailbreak Detection for (Almost) Free!
por: Chen, Guorui, et al.
Publicado: (2025) -
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025) -
PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
por: Shen, Guobin, et al.
Publicado: (2025) -
Proactive defense against LLM Jailbreak
por: Zhao, Weiliang, et al.
Publicado: (2025) -
Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking
por: Zhang, Junke, et al.
Publicado: (2026)