Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Shengyun, Wang, Xinpeng, Zhai, Guangyao, Navab, Nassir, Plank, Barbara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912289343406080
author Si, Shengyun
Wang, Xinpeng
Zhai, Guangyao
Navab, Nassir
Plank, Barbara
author_facet Si, Shengyun
Wang, Xinpeng
Zhai, Guangyao
Navab, Nassir
Plank, Barbara
contents Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17882
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
Si, Shengyun
Wang, Xinpeng
Zhai, Guangyao
Navab, Nassir
Plank, Barbara
Computation and Language
Artificial Intelligence
Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection.
title Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.17882