GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Yue, Zhai, Shengfang, Du, Mingzhe, Chen, Yulin, Cao, Tri, Gao, Hongcheng, Wang, Cheng, Li, Xinfeng, Wang, Kun, Fang, Junfeng, Zhang, Jiaheng, Hooi, Bryan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908367142780928
author Liu, Yue
Zhai, Shengfang
Du, Mingzhe
Chen, Yulin
Cao, Tri
Gao, Hongcheng
Wang, Cheng
Li, Xinfeng
Wang, Kun
Fang, Junfeng
Zhang, Jiaheng
Hooi, Bryan
author_facet Liu, Yue
Zhai, Shengfang
Du, Mingzhe
Chen, Yulin
Cao, Tri
Gao, Hongcheng
Wang, Cheng
Li, Xinfeng
Wang, Kun
Fang, Junfeng
Zhang, Jiaheng
Hooi, Bryan
contents To enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL. First, we construct GuardReasoner-VLTrain, a reasoning corpus with 123K samples and 631K reasoning steps, spanning text, image, and text-image inputs. Then, based on it, we cold-start our model's reasoning ability via SFT. In addition, we further enhance reasoning regarding moderation through online RL. Concretely, to enhance diversity and difficulty of samples, we conduct rejection sampling followed by data augmentation via the proposed safety-aware data concatenation. Besides, we use a dynamic clipping parameter to encourage exploration in early stages and exploitation in later stages. To balance performance and token efficiency, we design a length-aware safety reward that integrates accuracy, format, and token cost. Extensive experiments demonstrate the superiority of our model. Remarkably, it surpasses the runner-up by 19.27% F1 score on average. We release data, code, and models (3B/7B) of GuardReasoner-VL at https://github.com/yueliu1999/GuardReasoner-VL/
format Preprint
id arxiv_https___arxiv_org_abs_2505_11049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
Liu, Yue
Zhai, Shengfang
Du, Mingzhe
Chen, Yulin
Cao, Tri
Gao, Hongcheng
Wang, Cheng
Li, Xinfeng
Wang, Kun
Fang, Junfeng
Zhang, Jiaheng
Hooi, Bryan
Artificial Intelligence
Cryptography and Security
To enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL. First, we construct GuardReasoner-VLTrain, a reasoning corpus with 123K samples and 631K reasoning steps, spanning text, image, and text-image inputs. Then, based on it, we cold-start our model's reasoning ability via SFT. In addition, we further enhance reasoning regarding moderation through online RL. Concretely, to enhance diversity and difficulty of samples, we conduct rejection sampling followed by data augmentation via the proposed safety-aware data concatenation. Besides, we use a dynamic clipping parameter to encourage exploration in early stages and exploitation in later stages. To balance performance and token efficiency, we design a length-aware safety reward that integrates accuracy, format, and token cost. Extensive experiments demonstrate the superiority of our model. Remarkably, it surpasses the runner-up by 19.27% F1 score on average. We release data, code, and models (3B/7B) of GuardReasoner-VL at https://github.com/yueliu1999/GuardReasoner-VL/
title GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2505.11049