Steering Over-refusals Towards Safety in Retrieval Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maskey, Utsav, Dras, Mark, Naseem, Usman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915548576612352
author Maskey, Utsav
Dras, Mark
Naseem, Usman
author_facet Maskey, Utsav
Dras, Mark
Naseem, Usman
contents Safety alignment in large language models (LLMs) induces over-refusals -- where LLMs decline benign requests due to aggressive safety filters. We analyze this phenomenon in retrieval-augmented generation (RAG), where both the query intent and retrieved context properties influence refusal behavior. We construct RagRefuse, a domain-stratified benchmark spanning medical, chemical, and open domains, pairing benign and harmful queries with controlled context contamination patterns and sizes. Our analysis shows that context arrangement / contamination, domain of query and context, and harmful-text density trigger refusals even on benign queries, with effects depending on model-specific alignment choices. To mitigate over-refusals, we introduce \textsc{SafeRAG-Steering}, a model-centric embedding intervention that steers the embedding regions towards the confirmed safe, non-refusing output regions at inference time. This reduces over-refusals in contaminated RAG pipelines while preserving legitimate refusals.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10452
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steering Over-refusals Towards Safety in Retrieval Augmented Generation
Maskey, Utsav
Dras, Mark
Naseem, Usman
Computation and Language
Safety alignment in large language models (LLMs) induces over-refusals -- where LLMs decline benign requests due to aggressive safety filters. We analyze this phenomenon in retrieval-augmented generation (RAG), where both the query intent and retrieved context properties influence refusal behavior. We construct RagRefuse, a domain-stratified benchmark spanning medical, chemical, and open domains, pairing benign and harmful queries with controlled context contamination patterns and sizes. Our analysis shows that context arrangement / contamination, domain of query and context, and harmful-text density trigger refusals even on benign queries, with effects depending on model-specific alignment choices. To mitigate over-refusals, we introduce \textsc{SafeRAG-Steering}, a model-centric embedding intervention that steers the embedding regions towards the confirmed safe, non-refusing output regions at inference time. This reduces over-refusals in contaminated RAG pipelines while preserving legitimate refusals.
title Steering Over-refusals Towards Safety in Retrieval Augmented Generation
topic Computation and Language
url https://arxiv.org/abs/2510.10452