Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912289343406080 |
|---|---|
| author | Si, Shengyun Wang, Xinpeng Zhai, Guangyao Navab, Nassir Plank, Barbara |
| author_facet | Si, Shengyun Wang, Xinpeng Zhai, Guangyao Navab, Nassir Plank, Barbara |
| contents | Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_17882 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior Si, Shengyun Wang, Xinpeng Zhai, Guangyao Navab, Nassir Plank, Barbara Computation and Language Artificial Intelligence Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection. |
| title | Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2503.17882 |