SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Tinghao, Qi, Xiangyu, Zeng, Yi, Huang, Yangsibo, Sehwag, Udari Madhushani, Huang, Kaixuan, He, Luxi, Wei, Boyi, Li, Dacheng, Sheng, Ying, Jia, Ruoxi, Li, Bo, Li, Kai, Chen, Danqi, Henderson, Peter, Mittal, Prateek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
AI Risk Management Should Incorporate Both Safety and Security
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
The Model Hears You: Audio Language Model Deployments Should Consider the Principle of Least Privilege
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
Fantastic Copyrighted Beasts and How (Not) to Generate Them
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Can LLMs be Scammed? A Baseline Measurement Study
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
An Adversarial Perspective on Machine Unlearning for AI Safety
by: Łucki, Jakub, et al.
Published: (2024)
by: Łucki, Jakub, et al.
Published: (2024)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
by: Cook, Thomas, et al.
Published: (2025)
by: Cook, Thomas, et al.
Published: (2025)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
Evaluating Copyright Takedown Methods for Language Models
by: Wei, Boyi, et al.
Published: (2024)
by: Wei, Boyi, et al.
Published: (2024)
Efficient Data Shapley for Weighted Nearest Neighbor Algorithms
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
by: Ren, Kaixuan, et al.
Published: (2025)
by: Ren, Kaixuan, et al.
Published: (2025)
AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
by: Karthikeyan, Harish, et al.
Published: (2025)
by: Karthikeyan, Harish, et al.
Published: (2025)
Adapting to Evolving Adversaries with Regularized Continual Robust Training
by: Dai, Sihui, et al.
Published: (2025)
by: Dai, Sihui, et al.
Published: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
by: Panda, Ashwinee, et al.
Published: (2022)
by: Panda, Ashwinee, et al.
Published: (2022)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Averaging quadratically twisted $L$-values and their derivatives
by: Huang, Tinghao
Published: (2025)
by: Huang, Tinghao
Published: (2025)
Data Shapley in One Training Run
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026)
by: Pu, George, et al.
Published: (2026)
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
by: Zhang, Zhihao, et al.
Published: (2026)
by: Zhang, Zhihao, et al.
Published: (2026)
Capturing the Temporal Dependence of Training Data Influence
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
Certifiably Robust RAG against Retrieval Corruption
by: Xiang, Chong, et al.
Published: (2024)
by: Xiang, Chong, et al.
Published: (2024)
Does More Inference-Time Compute Really Help Robustness?
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
StructSR: Refuse Spurious Details in Real-World Image Super-Resolution
by: Li, Yachao, et al.
Published: (2025)
by: Li, Yachao, et al.
Published: (2025)
On Ramanujan Primes for Hecke-Maass Cusp Forms
by: Huang, Tinghao, et al.
Published: (2026)
by: Huang, Tinghao, et al.
Published: (2026)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
by: Lee, Michael S., et al.
Published: (2026)
by: Lee, Michael S., et al.
Published: (2026)
Statutory Construction and Interpretation for Artificial Intelligence
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
by: Xiao, Yuchen, et al.
Published: (2023)
by: Xiao, Yuchen, et al.
Published: (2023)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
by: Yuan, Youliang, et al.
Published: (2024)
by: Yuan, Youliang, et al.
Published: (2024)
Adaptive and Stratified Subsampling for High-Dimensional Robust Estimation
by: Mittal, Prateek, et al.
Published: (2024)
by: Mittal, Prateek, et al.
Published: (2024)
Similar Items
-
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
by: Wei, Boyi, et al.
Published: (2024) -
AI Risk Management Should Incorporate Both Safety and Security
by: Qi, Xiangyu, et al.
Published: (2024) -
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024) -
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025) -
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)