Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chhabra, Vishnu Kabir, Khalili, Mohammad Mahdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916675227484160
author Chhabra, Vishnu Kabir
Khalili, Mohammad Mahdi
author_facet Chhabra, Vishnu Kabir
Khalili, Mohammad Mahdi
contents The rapid growth of large language models has spurred significant interest in model compression as a means to enhance their accessibility and practicality. While extensive research has explored model compression through the lens of safety, findings suggest that safety-aligned models often lose elements of trustworthiness post-compression. Simultaneously, the field of mechanistic interpretability has gained traction, with notable discoveries, such as the identification of a single direction in the residual stream mediating refusal behaviors across diverse model architectures. In this work, we investigate the safety of compressed models by examining the mechanisms of refusal, adopting a novel interpretability-driven perspective to evaluate model safety. Furthermore, leveraging insights from our interpretability analysis, we propose a lightweight, computationally efficient method to enhance the safety of compressed models without compromising their performance or utility.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04215
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
Chhabra, Vishnu Kabir
Khalili, Mohammad Mahdi
Computation and Language
Artificial Intelligence
The rapid growth of large language models has spurred significant interest in model compression as a means to enhance their accessibility and practicality. While extensive research has explored model compression through the lens of safety, findings suggest that safety-aligned models often lose elements of trustworthiness post-compression. Simultaneously, the field of mechanistic interpretability has gained traction, with notable discoveries, such as the identification of a single direction in the residual stream mediating refusal behaviors across diverse model architectures. In this work, we investigate the safety of compressed models by examining the mechanisms of refusal, adopting a novel interpretability-driven perspective to evaluate model safety. Furthermore, leveraging insights from our interpretability analysis, we propose a lightweight, computationally efficient method to enhance the safety of compressed models without compromising their performance or utility.
title Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.04215