A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Xingjun, Wang, Yixu, Xu, Hengyuan, Wu, Yutao, Ding, Yifan, Zhao, Yunhan, Wang, Zilong, Hua, Jiabin, Wen, Ming, Liu, Jianan, Duan, Ranjie, Gao, Yifeng, Tan, Yingshui, Chen, Yunhao, Xue, Hui, Wang, Xin, Cheng, Wei, Chen, Jingjing, Wu, Zuxuan, Li, Bo, Jiang, Yu-Gang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908770501656576
author Ma, Xingjun
Wang, Yixu
Xu, Hengyuan
Wu, Yutao
Ding, Yifan
Zhao, Yunhan
Wang, Zilong
Hua, Jiabin
Wen, Ming
Liu, Jianan
Duan, Ranjie
Gao, Yifeng
Tan, Yingshui
Chen, Yunhao
Xue, Hui
Wang, Xin
Cheng, Wei
Chen, Jingjing
Wu, Zuxuan
Li, Bo
Jiang, Yu-Gang
author_facet Ma, Xingjun
Wang, Yixu
Xu, Hengyuan
Wu, Yutao
Ding, Yifan
Zhao, Yunhan
Wang, Zilong
Hua, Jiabin
Wen, Ming
Liu, Jianan
Duan, Ranjie
Gao, Yifeng
Tan, Yingshui
Chen, Yunhao
Xue, Hui
Wang, Xin
Cheng, Wei
Chen, Jingjing
Wu, Zuxuan
Li, Bo
Jiang, Yu-Gang
contents The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and vision, yet whether these advances translate into comparable improvements in safety remains unclear, partly due to fragmented evaluations that focus on isolated modalities or threat models. In this report, we present an integrated safety evaluation of six frontier models--GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5--assessing each across language, vision-language, and image generation using a unified protocol that combines benchmark, adversarial, multilingual, and compliance evaluations. By aggregating results into safety leaderboards and model profiles, we reveal a highly uneven safety landscape: while GPT-5.2 demonstrates consistently strong and balanced performance, other models exhibit clear trade-offs across benchmark safety, adversarial robustness, multilingual generalization, and regulatory compliance. Despite strong results under standard benchmarks, all models remain highly vulnerable under adversarial testing, with worst-case safety rates dropping below 6%. Text-to-image models show slightly stronger alignment in regulated visual risk categories, yet remain fragile when faced with adversarial or semantically ambiguous prompts. Overall, these findings highlight that safety in frontier models is inherently multidimensional--shaped by modality, language, and evaluation design--underscoring the need for standardized, holistic safety assessments to better reflect real-world risk and guide responsible deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10527
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
Ma, Xingjun
Wang, Yixu
Xu, Hengyuan
Wu, Yutao
Ding, Yifan
Zhao, Yunhan
Wang, Zilong
Hua, Jiabin
Wen, Ming
Liu, Jianan
Duan, Ranjie
Gao, Yifeng
Tan, Yingshui
Chen, Yunhao
Xue, Hui
Wang, Xin
Cheng, Wei
Chen, Jingjing
Wu, Zuxuan
Li, Bo
Jiang, Yu-Gang
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and vision, yet whether these advances translate into comparable improvements in safety remains unclear, partly due to fragmented evaluations that focus on isolated modalities or threat models. In this report, we present an integrated safety evaluation of six frontier models--GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5--assessing each across language, vision-language, and image generation using a unified protocol that combines benchmark, adversarial, multilingual, and compliance evaluations. By aggregating results into safety leaderboards and model profiles, we reveal a highly uneven safety landscape: while GPT-5.2 demonstrates consistently strong and balanced performance, other models exhibit clear trade-offs across benchmark safety, adversarial robustness, multilingual generalization, and regulatory compliance. Despite strong results under standard benchmarks, all models remain highly vulnerable under adversarial testing, with worst-case safety rates dropping below 6%. Text-to-image models show slightly stronger alignment in regulated visual risk categories, yet remain fragile when faced with adversarial or semantically ambiguous prompts. Overall, these findings highlight that safety in frontier models is inherently multidimensional--shaped by modality, language, and evaluation design--underscoring the need for standardized, holistic safety assessments to better reflect real-world risk and guide responsible deployment.
title A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2601.10527