A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xingjun, Wang, Yixu, Xu, Hengyuan, Wu, Yutao, Ding, Yifan, Zhao, Yunhan, Wang, Zilong, Hua, Jiabin, Wen, Ming, Liu, Jianan, Duan, Ranjie, Gao, Yifeng, Tan, Yingshui, Chen, Yunhao, Xue, Hui, Wang, Xin, Cheng, Wei, Chen, Jingjing, Wu, Zuxuan, Li, Bo, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Qwen2.5-VL Technical Report
by: Bai, Shuai, et al.
Published: (2025)
by: Bai, Shuai, et al.
Published: (2025)
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
by: Jegham, Nidhal, et al.
Published: (2025)
by: Jegham, Nidhal, et al.
Published: (2025)
Qwen3-VL Technical Report
by: Bai, Shuai, et al.
Published: (2025)
by: Bai, Shuai, et al.
Published: (2025)
Seedream 3.0 Technical Report
by: Gao, Yu, et al.
Published: (2025)
by: Gao, Yu, et al.
Published: (2025)
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
by: Li, Mingxin, et al.
Published: (2026)
by: Li, Mingxin, et al.
Published: (2026)
Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale
by: Tang, Shengji, et al.
Published: (2026)
by: Tang, Shengji, et al.
Published: (2026)
Gemini Pro Defeated by GPT-4V: Evidence from Education
by: Lee, Gyeong-Geon, et al.
Published: (2023)
by: Lee, Gyeong-Geon, et al.
Published: (2023)
Seedream 4.0: Toward Next-generation Multimodal Image Generation
by: Seedream, Team, et al.
Published: (2025)
by: Seedream, Team, et al.
Published: (2025)
Tuning Qwen2.5-VL to Improve Its Web Interaction Skills
by: Yakovleva, Alexandra, et al.
Published: (2026)
by: Yakovleva, Alexandra, et al.
Published: (2026)
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
by: Shen, Jinjie, et al.
Published: (2026)
by: Shen, Jinjie, et al.
Published: (2026)
ModelLock: Locking Your Model With a Spell
by: Gao, Yifeng, et al.
Published: (2024)
by: Gao, Yifeng, et al.
Published: (2024)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and Reading
by: Wu, Yutao, et al.
Published: (2025)
by: Wu, Yutao, et al.
Published: (2025)
Towards Context-Invariant Safety Alignment for Large Language Models
by: Wang, Yixu, et al.
Published: (2026)
by: Wang, Yixu, et al.
Published: (2026)
StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Banana100: Breaking NR-IQA Metrics by 100 Iterative Image Replications with Nano Banana Pro
by: Tang, Kenan, et al.
Published: (2026)
by: Tang, Kenan, et al.
Published: (2026)
HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models
by: Chen, Zixing, et al.
Published: (2026)
by: Chen, Zixing, et al.
Published: (2026)
Pro‐Vitamin A Biofortified Cavendish Banana: Trait Stability in the Field
by: Jimmy M. Tindamanyire, et al.
Published: (2026)
by: Jimmy M. Tindamanyire, et al.
Published: (2026)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
by: Chen, Yunhao, et al.
Published: (2025)
by: Chen, Yunhao, et al.
Published: (2025)
Qwen3.5-Omni Technical Report
by: Qwen Team
Published: (2026)
by: Qwen Team
Published: (2026)
Qwen3-VL-Seg: Unlocking Open-World Referring Segmentation with Vision-Language Grounding
by: Yao, Yuan, et al.
Published: (2026)
by: Yao, Yuan, et al.
Published: (2026)
Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
by: Wang, Weizhi, et al.
Published: (2025)
by: Wang, Weizhi, et al.
Published: (2025)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
by: Si, Shengyu, et al.
Published: (2026)
by: Si, Shengyu, et al.
Published: (2026)
ProMist-5K: A Comprehensive Dataset for Digital Emulation of Cinematic Pro-Mist Filter Effects
by: Lei, Yingtie, et al.
Published: (2026)
by: Lei, Yingtie, et al.
Published: (2026)
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
HoneypotNet: Backdoor Attacks Against Model Extraction
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
by: Gao, Yifeng, et al.
Published: (2025)
by: Gao, Yifeng, et al.
Published: (2025)
Qwen2.5-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
by: Xu, Pusheng, et al.
Published: (2025)
by: Xu, Pusheng, et al.
Published: (2025)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
by: Li, Juncheng, et al.
Published: (2025)
by: Li, Juncheng, et al.
Published: (2025)
Automated data extraction for systematic reviews using GPT‐5.2 and Google Gemini Pro 3: A dual‐large language model approach in orthopaedic research
by: Prushoth Vivekanantha, et al.
Published: (2026)
by: Prushoth Vivekanantha, et al.
Published: (2026)
Pro-Tensor Network
by: Yue, Gen, et al.
Published: (2026)
by: Yue, Gen, et al.
Published: (2026)
Qwen3-ASR Technical Report
by: Shi, Xian, et al.
Published: (2026)
by: Shi, Xian, et al.
Published: (2026)
SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model
by: Cazzaniga, Luca
Published: (2026)
by: Cazzaniga, Luca
Published: (2026)
Seed1.5-VL Technical Report
by: Guo, Dong, et al.
Published: (2025)
by: Guo, Dong, et al.
Published: (2025)
GaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enhanced Quality
by: Yi, Taoran, et al.
Published: (2024)
by: Yi, Taoran, et al.
Published: (2024)
Qwen3-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
Similar Items
-
Qwen2.5-VL Technical Report
by: Bai, Shuai, et al.
Published: (2025) -
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
by: Jegham, Nidhal, et al.
Published: (2025) -
Qwen3-VL Technical Report
by: Bai, Shuai, et al.
Published: (2025) -
Seedream 3.0 Technical Report
by: Gao, Yu, et al.
Published: (2025) -
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
by: Feng, Yunhao, et al.
Published: (2026)