Safe Vision-Language Models via Unsafe Weights Manipulation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: D'Incà, Moreno, Peruzzo, Elia, Xu, Xingqian, Shi, Humphrey, Sebe, Nicu, Mancini, Massimiliano
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911366916341760
author D'Incà, Moreno
Peruzzo, Elia
Xu, Xingqian
Shi, Humphrey
Sebe, Nicu
Mancini, Massimiliano
author_facet D'Incà, Moreno
Peruzzo, Elia
Xu, Xingqian
Shi, Humphrey
Sebe, Nicu
Mancini, Massimiliano
contents Vision-language models (VLMs) often inherit the biases and unsafe associations present within their large-scale training dataset. While recent approaches mitigate unsafe behaviors, their evaluation focuses on how safe the model is on unsafe inputs, ignoring potential shortcomings on safe ones. In this paper, we first revise safety evaluation by introducing SafeGround, a new set of metrics that evaluate safety at different levels of granularity. With this metric, we uncover a surprising issue of training-based methods: they make the model less safe on safe inputs. From this finding, we take a different direction and explore whether it is possible to make a model safer without training, introducing Unsafe Weights Manipulation (UWM). UWM uses a calibration set of safe and unsafe instances to compare activations between safe and unsafe content, identifying the most important parameters for processing the latter. Their values are then manipulated via negation. Experiments show that UWM achieves the best tradeoff between safety and knowledge preservation, consistently improving VLMs on unsafe queries while outperforming even training-based state-of-the-art methods on safe ones.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11742
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Safe Vision-Language Models via Unsafe Weights Manipulation
D'Incà, Moreno
Peruzzo, Elia
Xu, Xingqian
Shi, Humphrey
Sebe, Nicu
Mancini, Massimiliano
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-language models (VLMs) often inherit the biases and unsafe associations present within their large-scale training dataset. While recent approaches mitigate unsafe behaviors, their evaluation focuses on how safe the model is on unsafe inputs, ignoring potential shortcomings on safe ones. In this paper, we first revise safety evaluation by introducing SafeGround, a new set of metrics that evaluate safety at different levels of granularity. With this metric, we uncover a surprising issue of training-based methods: they make the model less safe on safe inputs. From this finding, we take a different direction and explore whether it is possible to make a model safer without training, introducing Unsafe Weights Manipulation (UWM). UWM uses a calibration set of safe and unsafe instances to compare activations between safe and unsafe content, identifying the most important parameters for processing the latter. Their values are then manipulated via negation. Experiments show that UWM achieves the best tradeoff between safety and knowledge preservation, consistently improving VLMs on unsafe queries while outperforming even training-based state-of-the-art methods on safe ones.
title Safe Vision-Language Models via Unsafe Weights Manipulation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.11742