Mitigating Many-Shot Jailbreaking
Fuente:
arXiv
Guardado en:
| Autores principales: | Ackerman, Christopher M., Panickssery, Nina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024)
por: Peng, Benji, et al.
Publicado: (2024)
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025)
por: Nikolić, Kristina, et al.
Publicado: (2025)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
por: Ackerman, Christopher, et al.
Publicado: (2024)
por: Ackerman, Christopher, et al.
Publicado: (2024)
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
por: Hossain, Ismail, et al.
Publicado: (2026)
por: Hossain, Ismail, et al.
Publicado: (2026)
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
por: Chen, Kejia, et al.
Publicado: (2026)
por: Chen, Kejia, et al.
Publicado: (2026)
CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs
por: Arachchige, Shashie Dilhara Batan, et al.
Publicado: (2025)
por: Arachchige, Shashie Dilhara Batan, et al.
Publicado: (2025)
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
por: Zheng, Xiaosen, et al.
Publicado: (2024)
por: Zheng, Xiaosen, et al.
Publicado: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
por: Chen, Zhuowei, et al.
Publicado: (2025)
por: Chen, Zhuowei, et al.
Publicado: (2025)
A Causal Perspective for Enhancing Jailbreak Attack and Defense
por: Pan, Licheng, et al.
Publicado: (2026)
por: Pan, Licheng, et al.
Publicado: (2026)
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks
por: Nurlanov, Zhakshylyk, et al.
Publicado: (2026)
por: Nurlanov, Zhakshylyk, et al.
Publicado: (2026)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
por: Lin, Shuyi, et al.
Publicado: (2025)
por: Lin, Shuyi, et al.
Publicado: (2025)
Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs
por: Chen, Wenyu, et al.
Publicado: (2026)
por: Chen, Wenyu, et al.
Publicado: (2026)
Improved Large Language Model Jailbreak Detection via Pretrained Embeddings
por: Galinkin, Erick, et al.
Publicado: (2024)
por: Galinkin, Erick, et al.
Publicado: (2024)
Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
por: Li, Songze, et al.
Publicado: (2025)
por: Li, Songze, et al.
Publicado: (2025)
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
por: Wu, Yuanwei, et al.
Publicado: (2023)
por: Wu, Yuanwei, et al.
Publicado: (2023)
PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline
por: Rao, Akshaj Prashanth, et al.
Publicado: (2025)
por: Rao, Akshaj Prashanth, et al.
Publicado: (2025)
Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
por: Yang, Junxiao, et al.
Publicado: (2025)
por: Yang, Junxiao, et al.
Publicado: (2025)
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
por: Liu, Xiaogeng, et al.
Publicado: (2024)
por: Liu, Xiaogeng, et al.
Publicado: (2024)
Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks
por: Wang, Zi, et al.
Publicado: (2024)
por: Wang, Zi, et al.
Publicado: (2024)
MoJE: Mixture of Jailbreak Experts, Naive Tabular Classifiers as Guard for Prompt Attacks
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)
por: Cornacchia, Giandomenico, et al.
Publicado: (2024)
Jailbreaking in the Haystack
por: Shah, Rishi Rajesh, et al.
Publicado: (2025)
por: Shah, Rishi Rajesh, et al.
Publicado: (2025)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
por: Chu, Junjie, et al.
Publicado: (2024)
por: Chu, Junjie, et al.
Publicado: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
por: Ghosal, Soumya Suvra, et al.
Publicado: (2024)
por: Ghosal, Soumya Suvra, et al.
Publicado: (2024)
Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
por: Dong, Yingkai, et al.
Publicado: (2024)
por: Dong, Yingkai, et al.
Publicado: (2024)
Jailbreaking with Universal Multi-Prompts
por: Hsu, Yu-Ling, et al.
Publicado: (2025)
por: Hsu, Yu-Ling, et al.
Publicado: (2025)
Jailbreaking LLMs via Calibration
por: Lu, Yuxuan, et al.
Publicado: (2026)
por: Lu, Yuxuan, et al.
Publicado: (2026)
Mitigating the Structural Bias in Graph Adversarial Defenses
por: Fang, Junyuan, et al.
Publicado: (2025)
por: Fang, Junyuan, et al.
Publicado: (2025)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
por: Li, Nan, et al.
Publicado: (2024)
por: Li, Nan, et al.
Publicado: (2024)
Fair Finetuning Mitigates Distribution Inference Attacks
por: Naidu, Rakshit
Publicado: (2026)
por: Naidu, Rakshit
Publicado: (2026)
In-Context Unlearning: Language Models as Few Shot Unlearners
por: Pawelczyk, Martin, et al.
Publicado: (2023)
por: Pawelczyk, Martin, et al.
Publicado: (2023)
Optimal Zero-Shot Detector for Multi-Armed Attacks
por: Granese, Federica, et al.
Publicado: (2024)
por: Granese, Federica, et al.
Publicado: (2024)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
por: Ren, Zhiyao, et al.
Publicado: (2025)
por: Ren, Zhiyao, et al.
Publicado: (2025)
Adaptive PII Mitigation Framework for Large Language Models
por: Asthana, Shubhi, et al.
Publicado: (2025)
por: Asthana, Shubhi, et al.
Publicado: (2025)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
por: Min, Rui, et al.
Publicado: (2024)
por: Min, Rui, et al.
Publicado: (2024)
A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning
por: Gluch, Greg, et al.
Publicado: (2025)
por: Gluch, Greg, et al.
Publicado: (2025)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
por: Vyas, Sanyam, et al.
Publicado: (2024)
por: Vyas, Sanyam, et al.
Publicado: (2024)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
por: Han, Tingxu, et al.
Publicado: (2024)
por: Han, Tingxu, et al.
Publicado: (2024)
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
por: Zhu, Sicheng, et al.
Publicado: (2024)
por: Zhu, Sicheng, et al.
Publicado: (2024)
Rethinking How to Evaluate Language Model Jailbreak
por: Cai, Hongyu, et al.
Publicado: (2024)
por: Cai, Hongyu, et al.
Publicado: (2024)
Ejemplares similares
-
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
por: Peng, Benji, et al.
Publicado: (2024) -
The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
por: Nikolić, Kristina, et al.
Publicado: (2025) -
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
por: Ackerman, Christopher, et al.
Publicado: (2024) -
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring
por: Hossain, Ismail, et al.
Publicado: (2026) -
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
por: Chen, Kejia, et al.
Publicado: (2026)