Rapid Response: Mitigating LLM Jailbreaks with a Few Examples
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Alwin, Michael, Julian, Sleight, Henry, Perez, Ethan, Sharma, Mrinank |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024)
by: Hughes, John, et al.
Published: (2024)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024)
by: Wang, Tony T., et al.
Published: (2024)
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
by: Youstra, Jack, et al.
Published: (2025)
by: Youstra, Jack, et al.
Published: (2025)
Who's in Charge? Disempowerment Patterns in Real-World LLM Usage
by: Sharma, Mrinank, et al.
Published: (2026)
by: Sharma, Mrinank, et al.
Published: (2026)
Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning
by: Sel, Bilgehan, et al.
Published: (2026)
by: Sel, Bilgehan, et al.
Published: (2026)
Sparse Autoencoders are Capable LLM Jailbreak Mitigators
by: Assogba, Yannick, et al.
Published: (2026)
by: Assogba, Yannick, et al.
Published: (2026)
All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
by: Guo, Shiyuan, et al.
Published: (2025)
by: Guo, Shiyuan, et al.
Published: (2025)
Stress-Testing Model Specs Reveals Character Differences among Language Models
by: Zhang, Jifan, et al.
Published: (2025)
by: Zhang, Jifan, et al.
Published: (2025)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
by: Ji, Haoxuan, et al.
Published: (2024)
by: Ji, Haoxuan, et al.
Published: (2024)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)
by: Huang, Ruixuan, et al.
Published: (2025)
Proactive defense against LLM Jailbreak
by: Zhao, Weiliang, et al.
Published: (2025)
by: Zhao, Weiliang, et al.
Published: (2025)
Mitigating Jailbreaks with Intent-Aware LLMs
by: Yeo, Wei Jie, et al.
Published: (2025)
by: Yeo, Wei Jie, et al.
Published: (2025)
Looking Inward: Language Models Can Learn About Themselves by Introspection
by: Binder, Felix J, et al.
Published: (2024)
by: Binder, Felix J, et al.
Published: (2024)
APE: Active Learning-based Tooling for Finding Informative Few-shot Examples for LLM-based Entity Matching
by: Qian, Kun, et al.
Published: (2024)
by: Qian, Kun, et al.
Published: (2024)
Believe It or Not: How Deeply do LLMs Believe Implanted Facts?
by: Slocum, Stewart, et al.
Published: (2025)
by: Slocum, Stewart, et al.
Published: (2025)
xJailbreak: Representation Space Guided Reinforcement Learning for Interpretable LLM Jailbreaking
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
Few-Shot Fairness: Unveiling LLM's Potential for Fairness-Aware Classification
by: Chhikara, Garima, et al.
Published: (2024)
by: Chhikara, Garima, et al.
Published: (2024)
Designing Informative Metrics for Few-Shot Example Selection
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Subtoxic Questions: Dive Into Attitude Change of LLM's Response in Jailbreak Attempts
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
GistScore: Learning Better Representations for In-Context Example Selection with Gist Bottlenecks
by: Gupta, Shivanshu, et al.
Published: (2023)
by: Gupta, Shivanshu, et al.
Published: (2023)
Jailbreaking to Jailbreak
by: Kritz, Jeremy, et al.
Published: (2025)
by: Kritz, Jeremy, et al.
Published: (2025)
Refactoring Programs Using Large Language Models with Few-Shot Examples
by: Shirafuji, Atsushi, et al.
Published: (2023)
by: Shirafuji, Atsushi, et al.
Published: (2023)
Learning New Tasks from a Few Examples with Soft-Label Prototypes
by: Singh, Avyav Kumar, et al.
Published: (2022)
by: Singh, Avyav Kumar, et al.
Published: (2022)
Inverse Scaling in Test-Time Compute
by: Gema, Aryo Pradipta, et al.
Published: (2025)
by: Gema, Aryo Pradipta, et al.
Published: (2025)
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
by: Yao, Jia-Yu, et al.
Published: (2023)
by: Yao, Jia-Yu, et al.
Published: (2023)
Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak Attacks
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
Few-Shot Multilingual Open-Domain QA from 5 Examples
by: Jiang, Fan, et al.
Published: (2025)
by: Jiang, Fan, et al.
Published: (2025)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
by: Fu, Tingchen, et al.
Published: (2024)
by: Fu, Tingchen, et al.
Published: (2024)
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
by: Shen, Guangyu, et al.
Published: (2024)
by: Shen, Guangyu, et al.
Published: (2024)
Not the Example, but the Process: How Self-Generated Examples Enhance LLM Reasoning
by: Gwak, Daehoon, et al.
Published: (2026)
by: Gwak, Daehoon, et al.
Published: (2026)
Factors That Support Grounded Responses in LLM Conversations: A Rapid Review
by: Iwashima, Gabriele Cesar, et al.
Published: (2025)
by: Iwashima, Gabriele Cesar, et al.
Published: (2025)
The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models
by: Yoshida, Lui
Published: (2024)
by: Yoshida, Lui
Published: (2024)
Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
by: Sharma, Mrinank, et al.
Published: (2025)
by: Sharma, Mrinank, et al.
Published: (2025)
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
by: Ouyang, Yang, et al.
Published: (2025)
by: Ouyang, Yang, et al.
Published: (2025)
Similar Items
-
Best-of-N Jailbreaking
by: Hughes, John, et al.
Published: (2024) -
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
by: Wang, Tony T., et al.
Published: (2024) -
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
by: Schaeffer, Rylan, et al.
Published: (2024) -
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
by: Wen, Jiaxin, et al.
Published: (2024) -
Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
by: Youstra, Jack, et al.
Published: (2025)