RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muhamed, Aashiq, Ribeiro, Leonardo F. R., Dreyer, Markus, Smith, Virginia, Diab, Mona T.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909839774449664
author Muhamed, Aashiq
Ribeiro, Leonardo F. R.
Dreyer, Markus
Smith, Virginia
Diab, Mona T.
author_facet Muhamed, Aashiq
Ribeiro, Leonardo F. R.
Dreyer, Markus
Smith, Virginia
Diab, Mona T.
contents The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-scale study reveals that even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks, while exhibiting either dangerous overconfidence or overcaution. Static benchmarks fail to reliably evaluate this capability, as models exploit dataset-specific artifacts and memorize test instances. We introduce RefusalBench, a generative methodology that programmatically creates diagnostic test cases through controlled linguistic perturbation. Our framework employs 176 distinct perturbation strategies across six categories of informational uncertainty and three intensity levels. Evaluation of over 30 models uncovers systematic failure patterns: refusal comprises separable detection and categorization skills, and neither scale nor extended reasoning improves performance. We find that selective refusal is a trainable, alignment-sensitive capability, offering a clear path for improvement. We release two benchmarks -- RefusalBench-NQ (single document) and RefusalBench-GaRAGe (multi-document) -- and our complete generation framework to enable continued, dynamic evaluation of this critical capability.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
Muhamed, Aashiq
Ribeiro, Leonardo F. R.
Dreyer, Markus
Smith, Virginia
Diab, Mona T.
Computation and Language
Artificial Intelligence
Machine Learning
The ability of language models in RAG systems to selectively refuse to answer based on flawed context is critical for safety, yet remains a significant failure point. Our large-scale study reveals that even frontier models struggle in this setting, with refusal accuracy dropping below 50% on multi-document tasks, while exhibiting either dangerous overconfidence or overcaution. Static benchmarks fail to reliably evaluate this capability, as models exploit dataset-specific artifacts and memorize test instances. We introduce RefusalBench, a generative methodology that programmatically creates diagnostic test cases through controlled linguistic perturbation. Our framework employs 176 distinct perturbation strategies across six categories of informational uncertainty and three intensity levels. Evaluation of over 30 models uncovers systematic failure patterns: refusal comprises separable detection and categorization skills, and neither scale nor extended reasoning improves performance. We find that selective refusal is a trainable, alignment-sensitive capability, offering a clear path for improvement. We release two benchmarks -- RefusalBench-NQ (single document) and RefusalBench-GaRAGe (multi-document) -- and our complete generation framework to enable continued, dynamic evaluation of this critical capability.
title RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.10390