CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Xiangyu, Liu, Jiaxu, Chen, Zhen, Hu, Jinwei, Dong, Yi, Huang, Xiaowei, Ruan, Wenjie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909547718770688
author Yin, Xiangyu
Liu, Jiaxu
Chen, Zhen
Hu, Jinwei
Dong, Yi
Huang, Xiaowei
Ruan, Wenjie
author_facet Yin, Xiangyu
Liu, Jiaxu
Chen, Zhen
Hu, Jinwei
Dong, Yi
Huang, Xiaowei
Ruan, Wenjie
contents Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailbreak attacks. First, we proposed a novel distance metric to quantify semantic discrepancies between malicious and intended responses, capturing subtle differences often overlooked by conventional cosine similarity-based measures. Then, we devise a regressed certification approach that employs randomized smoothing to provide formal robustness guarantees against both adversarial and structural perturbations, even under black-box settings. Complementing this, our feature-space defence introduces noise distributions (e.g., Gaussian, Laplacian) into the latent embeddings to safeguard against both pixel-level and structure-level perturbations. Our results highlight the potential of a formally grounded, integrated strategy toward building more resilient and trustworthy VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models
Yin, Xiangyu
Liu, Jiaxu
Chen, Zhen
Hu, Jinwei
Dong, Yi
Huang, Xiaowei
Ruan, Wenjie
Computer Vision and Pattern Recognition
Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailbreak attacks. First, we proposed a novel distance metric to quantify semantic discrepancies between malicious and intended responses, capturing subtle differences often overlooked by conventional cosine similarity-based measures. Then, we devise a regressed certification approach that employs randomized smoothing to provide formal robustness guarantees against both adversarial and structural perturbations, even under black-box settings. Complementing this, our feature-space defence introduces noise distributions (e.g., Gaussian, Laplacian) into the latent embeddings to safeguard against both pixel-level and structure-level perturbations. Our results highlight the potential of a formally grounded, integrated strategy toward building more resilient and trustworthy VLMs.
title CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.10661