CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909547718770688 |
|---|---|
| author | Yin, Xiangyu Liu, Jiaxu Chen, Zhen Hu, Jinwei Dong, Yi Huang, Xiaowei Ruan, Wenjie |
| author_facet | Yin, Xiangyu Liu, Jiaxu Chen, Zhen Hu, Jinwei Dong, Yi Huang, Xiaowei Ruan, Wenjie |
| contents | Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailbreak attacks. First, we proposed a novel distance metric to quantify semantic discrepancies between malicious and intended responses, capturing subtle differences often overlooked by conventional cosine similarity-based measures. Then, we devise a regressed certification approach that employs randomized smoothing to provide formal robustness guarantees against both adversarial and structural perturbations, even under black-box settings. Complementing this, our feature-space defence introduces noise distributions (e.g., Gaussian, Laplacian) into the latent embeddings to safeguard against both pixel-level and structure-level perturbations. Our results highlight the potential of a formally grounded, integrated strategy toward building more resilient and trustworthy VLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10661 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models Yin, Xiangyu Liu, Jiaxu Chen, Zhen Hu, Jinwei Dong, Yi Huang, Xiaowei Ruan, Wenjie Computer Vision and Pattern Recognition Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailbreak attacks. First, we proposed a novel distance metric to quantify semantic discrepancies between malicious and intended responses, capturing subtle differences often overlooked by conventional cosine similarity-based measures. Then, we devise a regressed certification approach that employs randomized smoothing to provide formal robustness guarantees against both adversarial and structural perturbations, even under black-box settings. Complementing this, our feature-space defence introduces noise distributions (e.g., Gaussian, Laplacian) into the latent embeddings to safeguard against both pixel-level and structure-level perturbations. Our results highlight the potential of a formally grounded, integrated strategy toward building more resilient and trustworthy VLMs. |
| title | CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.10661 |