FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Jiajun, Mao, Jiageng, Qi, Ang, Yuan, Weiduo, Romanus, Alexander, Xia, Helen, Guizilini, Vitor Campagnolo, Wang, Yue
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917319126548480
author Xu, Jiajun
Mao, Jiageng
Qi, Ang
Yuan, Weiduo
Romanus, Alexander
Xia, Helen
Guizilini, Vitor Campagnolo
Wang, Yue
author_facet Xu, Jiajun
Mao, Jiageng
Qi, Ang
Yuan, Weiduo
Romanus, Alexander
Xia, Helen
Guizilini, Vitor Campagnolo
Wang, Yue
contents Vision Language Models (VLMs) are prone to errors, and identifying where these errors occur is critical for ensuring the reliability and safety of AI systems. In this paper, we propose an approach that automatically generates questions designed to deliberately induce incorrect responses from VLMs, thereby revealing their vulnerabilities. The core of this approach lies in fuzz testing and reinforcement finetuning: we transform a single input query into a large set of diverse variants through vision and language fuzzing. Based on the fuzzing outcomes, the question generator is further instructed by adversarial reinforcement fine-tuning to produce increasingly challenging queries that trigger model failures. With this approach, we can consistently drive down a target VLM's answer accuracy -- for example, the accuracy of Qwen2.5-VL-32B on our generated questions drops from 86.58\% to 65.53\% in four RL iterations. Moreover, a fuzzing policy trained against a single target VLM transfers to multiple other VLMs, producing challenging queries that degrade their performance as well.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06600
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures
Xu, Jiajun
Mao, Jiageng
Qi, Ang
Yuan, Weiduo
Romanus, Alexander
Xia, Helen
Guizilini, Vitor Campagnolo
Wang, Yue
Machine Learning
Artificial Intelligence
Vision Language Models (VLMs) are prone to errors, and identifying where these errors occur is critical for ensuring the reliability and safety of AI systems. In this paper, we propose an approach that automatically generates questions designed to deliberately induce incorrect responses from VLMs, thereby revealing their vulnerabilities. The core of this approach lies in fuzz testing and reinforcement finetuning: we transform a single input query into a large set of diverse variants through vision and language fuzzing. Based on the fuzzing outcomes, the question generator is further instructed by adversarial reinforcement fine-tuning to produce increasingly challenging queries that trigger model failures. With this approach, we can consistently drive down a target VLM's answer accuracy -- for example, the accuracy of Qwen2.5-VL-32B on our generated questions drops from 86.58\% to 65.53\% in four RL iterations. Moreover, a fuzzing policy trained against a single target VLM transfers to multiple other VLMs, producing challenging queries that degrade their performance as well.
title FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.06600