Directional Embedding Smoothing for Robust Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ye, Liu, Jing, Koike-Akino, Toshiaki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908889378717696
author Wang, Ye
Liu, Jing
Koike-Akino, Toshiaki
author_facet Wang, Ye
Liu, Jing
Koike-Akino, Toshiaki
contents The safety and reliability of vision-language models (VLMs) are a crucial part of deploying trustworthy agentic AI systems. However, VLMs remain vulnerable to jailbreaking attacks that undermine their safety alignment to yield harmful outputs. In this work, we extend the Randomized Embedding Smoothing and Token Aggregation (RESTA) defense to VLMs and evaluate its performance against the JailBreakV-28K benchmark of multi-modal jailbreaking attacks. We find that RESTA is effective in reducing attack success rate over this diverse corpus of attacks, in particular, when employing directional embedding noise, where the injected noise is aligned with the original token embedding vectors. Our results demonstrate that RESTA can contribute to securing VLMs within agentic systems, as a lightweight, inference-time defense layer of an overall security framework.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15259
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Directional Embedding Smoothing for Robust Vision Language Models
Wang, Ye
Liu, Jing
Koike-Akino, Toshiaki
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
The safety and reliability of vision-language models (VLMs) are a crucial part of deploying trustworthy agentic AI systems. However, VLMs remain vulnerable to jailbreaking attacks that undermine their safety alignment to yield harmful outputs. In this work, we extend the Randomized Embedding Smoothing and Token Aggregation (RESTA) defense to VLMs and evaluate its performance against the JailBreakV-28K benchmark of multi-modal jailbreaking attacks. We find that RESTA is effective in reducing attack success rate over this diverse corpus of attacks, in particular, when employing directional embedding noise, where the injected noise is aligned with the original token embedding vectors. Our results demonstrate that RESTA can contribute to securing VLMs within agentic systems, as a lightweight, inference-time defense layer of an overall security framework.
title Directional Embedding Smoothing for Robust Vision Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2603.15259