Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914114040758272 |
|---|---|
| author | Pham, Anh Thalanki, Mihir Sun, Michael Chaloo, Aditya Gupta, Ankita Xia, Tian Mate, Aditya Nosakhare, Ehimwenma Srinivasan, Soundararajan |
| author_facet | Pham, Anh Thalanki, Mihir Sun, Michael Chaloo, Aditya Gupta, Ankita Xia, Tian Mate, Aditya Nosakhare, Ehimwenma Srinivasan, Soundararajan |
| contents | Large language models often lose previously aligned safety behaviors when fine-tuned on benign data, a phenomenon known as catastrophic forgetting. Prior work shows that adding random safety examples can mitigate this effect, but it remains unclear which examples are most effective. We propose a behavior-aware sampling framework that selects safety examples based on two complementary factors: instruction-response behavior (e.g., refusal versus compliance) and semantic diversity across harm categories. Systematic evaluation shows that this approach substantially reduces harmful outputs while maintaining helpfulness, achieving up to a 41% reduction in harmfulness with only 0.5% additional training data. These results highlight how targeted data selection can improve the safety and efficiency of fine-tuning at scale. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_21885 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning Pham, Anh Thalanki, Mihir Sun, Michael Chaloo, Aditya Gupta, Ankita Xia, Tian Mate, Aditya Nosakhare, Ehimwenma Srinivasan, Soundararajan Computation and Language Artificial Intelligence Large language models often lose previously aligned safety behaviors when fine-tuned on benign data, a phenomenon known as catastrophic forgetting. Prior work shows that adding random safety examples can mitigate this effect, but it remains unclear which examples are most effective. We propose a behavior-aware sampling framework that selects safety examples based on two complementary factors: instruction-response behavior (e.g., refusal versus compliance) and semantic diversity across harm categories. Systematic evaluation shows that this approach substantially reduces harmful outputs while maintaining helpfulness, achieving up to a 41% reduction in harmfulness with only 0.5% additional training data. These results highlight how targeted data selection can improve the safety and efficiency of fine-tuning at scale. |
| title | Preventing Catastrophic Forgetting: Behavior-Aware Sampling for Safer Language Model Fine-Tuning |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2510.21885 |