Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Fraser, Kathleen C., Dawkins, Hillary, Nejadgholi, Isar, Kiritchenko, Svetlana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
by: Guo, Rongchen, et al.
Published: (2024)
by: Guo, Rongchen, et al.
Published: (2024)
When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
The crime of being poor
by: Curto, Georgina, et al.
Published: (2023)
by: Curto, Georgina, et al.
Published: (2023)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
From Perceived Effectiveness to Measured Impact: Identity-Aware Evaluation of Automated Counter-Stereotypes
by: Kiritchenko, Svetlana, et al.
Published: (2025)
by: Kiritchenko, Svetlana, et al.
Published: (2025)
Gender-Neutral Machine Translation Strategies in Practice
by: Dawkins, Hillary, et al.
Published: (2025)
by: Dawkins, Hillary, et al.
Published: (2025)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025)
by: Curto, Georgina, et al.
Published: (2025)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images
by: Fraser, Kathleen C., et al.
Published: (2024)
by: Fraser, Kathleen C., et al.
Published: (2024)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
by: Vishnubhotla, Krishnapriya, et al.
Published: (2026)
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
by: Nejadgholi, Isar, et al.
Published: (2026)
by: Nejadgholi, Isar, et al.
Published: (2026)
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
by: Nejadgholi, Isar, et al.
Published: (2025)
by: Nejadgholi, Isar, et al.
Published: (2025)
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
by: Ghanadian, Hamideh, et al.
Published: (2024)
by: Ghanadian, Hamideh, et al.
Published: (2024)
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
by: Guo, Rongchen, et al.
Published: (2025)
by: Guo, Rongchen, et al.
Published: (2025)
Uncovering Bias in Large Vision-Language Models with Counterfactuals
by: Howard, Phillip, et al.
Published: (2024)
by: Howard, Phillip, et al.
Published: (2024)
Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation
by: Churina, Svetlana, et al.
Published: (2024)
by: Churina, Svetlana, et al.
Published: (2024)
Social and Ethical Risks Posed by General-Purpose LLMs for Settling Newcomers in Canada
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
by: Lai, Wen, et al.
Published: (2025)
by: Lai, Wen, et al.
Published: (2025)
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
by: Niazi, Ruhallah, et al.
Published: (2026)
by: Niazi, Ruhallah, et al.
Published: (2026)
Human-Centered AI Applications for Canada's Immigration Settlement Sector
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Enhancing BERT Fine-Tuning for Sentiment Analysis in Lower-Resourced Languages
by: Kubík, Jozef, et al.
Published: (2025)
by: Kubík, Jozef, et al.
Published: (2025)
Safety-Aware Fine-Tuning of Large Language Models
by: Choi, Hyeong Kyu, et al.
Published: (2024)
by: Choi, Hyeong Kyu, et al.
Published: (2024)
GR-SAP: Generative Replay for Safety Alignment Preservation during Fine-Tuning
by: Fang, Zhouxiang, et al.
Published: (2026)
by: Fang, Zhouxiang, et al.
Published: (2026)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning
by: Yi, Xin, et al.
Published: (2024)
by: Yi, Xin, et al.
Published: (2024)
Cross-Cultural Value Awareness in Large Vision-Language Models
by: Howard, Phillip, et al.
Published: (2026)
by: Howard, Phillip, et al.
Published: (2026)
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026)
by: Vennemeyer, Daniel, et al.
Published: (2026)
Evaluating Prompt-Based and Fine-Tuned Approaches to Czech Anaphora Resolution
by: Stano, Patrik, et al.
Published: (2025)
by: Stano, Patrik, et al.
Published: (2025)
EdgeFlowerTune: Evaluating Federated LLM Fine-Tuning Under Realistic Edge System Constraints
by: Geng, Jiaxiang, et al.
Published: (2026)
by: Geng, Jiaxiang, et al.
Published: (2026)
Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Are Smarter LLMs Safer? Exploring Safety-Reasoning Trade-offs in Prompting and Fine-Tuning
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning
by: Goel, Jyotin, et al.
Published: (2026)
by: Goel, Jyotin, et al.
Published: (2026)
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
by: Ponkshe, Kaustubh, et al.
Published: (2025)
by: Ponkshe, Kaustubh, et al.
Published: (2025)
Fine-Tuning and Evaluating Conversational AI for Agricultural Advisory
by: Singh, Sanyam, et al.
Published: (2026)
by: Singh, Sanyam, et al.
Published: (2026)
How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
by: Huang, Zixian, et al.
Published: (2026)
by: Huang, Zixian, et al.
Published: (2026)
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
Similar Items
-
Adaptable Moral Stances of Large Language Models on Sexist Content: Implications for Society and Gender Discourse
by: Guo, Rongchen, et al.
Published: (2024) -
When Detection Fails: The Power of Fine-Tuned Models to Generate Human-Like Social Media Text
by: Dawkins, Hillary, et al.
Published: (2025) -
Detecting AI-Generated Text: Factors Influencing Detectability with Current Methods
by: Fraser, Kathleen C., et al.
Published: (2024) -
The crime of being poor
by: Curto, Georgina, et al.
Published: (2023) -
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)