GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Yuxiao, Chen, Junchi, Jin, Zhenchao, Miao, Changtao, Yuan, Haojie, Chu, Qi, Gong, Tao, Yu, Nenghai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914171360116736
author Xiang, Yuxiao
Chen, Junchi
Jin, Zhenchao
Miao, Changtao
Yuan, Haojie
Chu, Qi
Gong, Tao
Yu, Nenghai
author_facet Xiang, Yuxiao
Chen, Junchi
Jin, Zhenchao
Miao, Changtao
Yuan, Haojie
Chu, Qi
Gong, Tao
Yu, Nenghai
contents Multimodal large reasoning models (MLRMs) are increasingly deployed for vision-language tasks that produce explicit intermediate rationales. However, reasoning traces can contain unsafe content even when the final answer is non-harmful, creating deployment risks. Existing multimodal safety guards primarily evaluate only the input question and the final answer, neglecting the intermediate reasoning process. This oversight allows undetected harm, such as biased inferences or policy-violating use of visual context, to emerge during reasoning. We introduce GuardTrace-VL, a vision-aware safety auditor that monitors the full Question-Thinking-Answer (QTA) pipeline via joint image-text analysis, enabling detection of unsafe content as it emerges in the reasoning stage. To support training and evaluation, we construct the GuardTrace dataset, which is generated through diverse prompting strategies and refined via a MLRM- and human-based voting and verification pipeline. Furthermore, we propose a three-stage progressive training scheme combined with the data refinement process, enabling the model to learn nuanced and context-dependent safety preferences according to different risk levels. On our proposed test set covering both in-domain and out-of-domain scenarios, GuardTrace-VL model achieves an F1 score of 93.1% on unsafe reasoning detection tasks, representing a 13.5% improvement in F1 score compared to the previous strongest multimodal safety defense methods. The codes will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
Xiang, Yuxiao
Chen, Junchi
Jin, Zhenchao
Miao, Changtao
Yuan, Haojie
Chu, Qi
Gong, Tao
Yu, Nenghai
Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
Multimodal large reasoning models (MLRMs) are increasingly deployed for vision-language tasks that produce explicit intermediate rationales. However, reasoning traces can contain unsafe content even when the final answer is non-harmful, creating deployment risks. Existing multimodal safety guards primarily evaluate only the input question and the final answer, neglecting the intermediate reasoning process. This oversight allows undetected harm, such as biased inferences or policy-violating use of visual context, to emerge during reasoning. We introduce GuardTrace-VL, a vision-aware safety auditor that monitors the full Question-Thinking-Answer (QTA) pipeline via joint image-text analysis, enabling detection of unsafe content as it emerges in the reasoning stage. To support training and evaluation, we construct the GuardTrace dataset, which is generated through diverse prompting strategies and refined via a MLRM- and human-based voting and verification pipeline. Furthermore, we propose a three-stage progressive training scheme combined with the data refinement process, enabling the model to learn nuanced and context-dependent safety preferences according to different risk levels. On our proposed test set covering both in-domain and out-of-domain scenarios, GuardTrace-VL model achieves an F1 score of 93.1% on unsafe reasoning detection tasks, representing a 13.5% improvement in F1 score compared to the previous strongest multimodal safety defense methods. The codes will be made publicly available.
title GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2511.20994