Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Yewon, Seol, Yumin, Kong, EunGyung, Jo, Minsoo, Kim, Taesup
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911518732320768
author Han, Yewon
Seol, Yumin
Kong, EunGyung
Jo, Minsoo
Kim, Taesup
author_facet Han, Yewon
Seol, Yumin
Kong, EunGyung
Jo, Minsoo
Kim, Taesup
contents Existing jailbreak defence frameworks for Large Vision-Language Models often suffer from a safety utility tradeoff, where strengthening safety inadvertently degrades performance on general visual-grounded reasoning tasks. In this work, we investigate whether safety and utility are inherently antagonistic objectives. We focus on a modality induced bias direction consistently observed across datasets, which arises from suboptimal coupling between the Large Language Model backbone and visual encoders. We further demonstrate that this direction undermines performance on both tasks. Leveraging this insight, we propose Two Birds, One Projection, an efficient inference time jailbreak defence that projects cross-modal features onto the null space of the identified bias direction to remove the corresponding components. Requiring only a single forward pass, our method effectively breaks the conventional tradeoff, simultaneously improving both safety and utility across diverse benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14825
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
Han, Yewon
Seol, Yumin
Kong, EunGyung
Jo, Minsoo
Kim, Taesup
Computer Vision and Pattern Recognition
Artificial Intelligence
Existing jailbreak defence frameworks for Large Vision-Language Models often suffer from a safety utility tradeoff, where strengthening safety inadvertently degrades performance on general visual-grounded reasoning tasks. In this work, we investigate whether safety and utility are inherently antagonistic objectives. We focus on a modality induced bias direction consistently observed across datasets, which arises from suboptimal coupling between the Large Language Model backbone and visual encoders. We further demonstrate that this direction undermines performance on both tasks. Leveraging this insight, we propose Two Birds, One Projection, an efficient inference time jailbreak defence that projects cross-modal features onto the null space of the identified bias direction to remove the corresponding components. Requiring only a single forward pass, our method effectively breaks the conventional tradeoff, simultaneously improving both safety and utility across diverse benchmarks.
title Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.14825