Endless Jailbreaks with Bijection Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Brian R. Y., Li, Maximilian, Tang, Leonard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Plentiful Jailbreaks with String Compositions
von: Huang, Brian R. Y.
Veröffentlicht: (2024)
von: Huang, Brian R. Y.
Veröffentlicht: (2024)
Endless Terminals: Scaling RL Environments for Terminal Agents
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2026)
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2026)
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025)
Inroads to a Structured Data Natural Language Bijection and the role of LLM annotation
von: Vente, Blake
Veröffentlicht: (2024)
von: Vente, Blake
Veröffentlicht: (2024)
Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings
von: Huang, Yue, et al.
Veröffentlicht: (2024)
von: Huang, Yue, et al.
Veröffentlicht: (2024)
Jailbreaking to Jailbreak
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
von: Kritz, Jeremy, et al.
Veröffentlicht: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
von: Murphy, Brendan, et al.
Veröffentlicht: (2025)
von: Murphy, Brendan, et al.
Veröffentlicht: (2025)
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
von: Zhou, Weikang, et al.
Veröffentlicht: (2024)
von: Zhou, Weikang, et al.
Veröffentlicht: (2024)
Verdict: A Library for Scaling Judge-Time Compute
von: Kalra, Nimit, et al.
Veröffentlicht: (2025)
von: Kalra, Nimit, et al.
Veröffentlicht: (2025)
Jailbreaking? One Step Is Enough!
von: Zheng, Weixiong, et al.
Veröffentlicht: (2024)
von: Zheng, Weixiong, et al.
Veröffentlicht: (2024)
GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication
von: Tang, Hua, et al.
Veröffentlicht: (2025)
von: Tang, Hua, et al.
Veröffentlicht: (2025)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
von: Zhou, Yukai, et al.
Veröffentlicht: (2024)
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Can Large Language Models Automatically Jailbreak GPT-4V?
von: Wu, Yuanwei, et al.
Veröffentlicht: (2024)
von: Wu, Yuanwei, et al.
Veröffentlicht: (2024)
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Jailbreaking Large Language Models with Morality Attacks
von: Su, Ying, et al.
Veröffentlicht: (2026)
von: Su, Ying, et al.
Veröffentlicht: (2026)
Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
von: Lin, Yuping, et al.
Veröffentlicht: (2024)
von: Lin, Yuping, et al.
Veröffentlicht: (2024)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
von: Ji, Haoxuan, et al.
Veröffentlicht: (2024)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
von: Liu, Songyang, et al.
Veröffentlicht: (2026)
Many-Turn Jailbreaking
von: Yang, Xianjun, et al.
Veröffentlicht: (2025)
von: Yang, Xianjun, et al.
Veröffentlicht: (2025)
SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
von: Feng, Mingqian, et al.
Veröffentlicht: (2026)
von: Feng, Mingqian, et al.
Veröffentlicht: (2026)
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
von: Lin, Liang, et al.
Veröffentlicht: (2025)
von: Lin, Liang, et al.
Veröffentlicht: (2025)
Weak-to-Strong Jailbreaking on Large Language Models
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
von: Huang, Yuting, et al.
Veröffentlicht: (2025)
von: Huang, Yuting, et al.
Veröffentlicht: (2025)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
von: Hawkins, John, et al.
Veröffentlicht: (2025)
von: Hawkins, John, et al.
Veröffentlicht: (2025)
Jailbreaking as a Reward Misspecification Problem
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
RoleBreak: Character Hallucination as a Jailbreak Attack in Role-Playing Systems
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
von: Tang, Yihong, et al.
Veröffentlicht: (2024)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
Towards a Characterization of Two-way Bijections in a Reversible Computational Model
von: Palazzo, Matteo, et al.
Veröffentlicht: (2025)
von: Palazzo, Matteo, et al.
Veröffentlicht: (2025)
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
von: Liu, Xuannan, et al.
Veröffentlicht: (2024)
Enhancing Jailbreak Attacks with Diversity Guidance
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
Persona Jailbreaking in Large Language Models
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Huang, Ruixuan, et al.
Veröffentlicht: (2025)
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
von: Li, Xiaoxia, et al.
Veröffentlicht: (2024)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
von: Zhang, Zhexin, et al.
Veröffentlicht: (2023)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
von: Beetham, James, et al.
Veröffentlicht: (2024)
von: Beetham, James, et al.
Veröffentlicht: (2024)
Dialogue Injection Attack: Jailbreaking LLMs through Context Manipulation
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
von: Meng, Wenlong, et al.
Veröffentlicht: (2025)
Structured Semantic Cloaking for Jailbreak Attacks on Large Language Models
von: Sun, Xiaobing, et al.
Veröffentlicht: (2026)
von: Sun, Xiaobing, et al.
Veröffentlicht: (2026)
Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models
von: Li, Jiahui, et al.
Veröffentlicht: (2024)
von: Li, Jiahui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Plentiful Jailbreaks with String Compositions
von: Huang, Brian R. Y.
Veröffentlicht: (2024) -
Endless Terminals: Scaling RL Environments for Terminal Agents
von: Gandhi, Kanishk, et al.
Veröffentlicht: (2026) -
LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
von: Fu, Jiyuan, et al.
Veröffentlicht: (2025) -
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
von: Lee, Sunbowen, et al.
Veröffentlicht: (2025) -
Inroads to a Structured Data Natural Language Bijection and the role of LLM annotation
von: Vente, Blake
Veröffentlicht: (2024)