_version_ 1866917109959753728
author Bai, Shuai
Cai, Yuxuan
Chen, Ruizhe
Chen, Keqin
Chen, Xionghui
Cheng, Zesen
Deng, Lianghao
Ding, Wei
Gao, Chang
Ge, Chunjiang
Ge, Wenbin
Guo, Zhifang
Huang, Qidong
Huang, Jie
Huang, Fei
Hui, Binyuan
Jiang, Shutong
Li, Zhaohai
Li, Mingsheng
Li, Mei
Li, Kaixin
Lin, Zicheng
Lin, Junyang
Liu, Xuejing
Liu, Jiawei
Liu, Chenglong
Liu, Yang
Liu, Dayiheng
Liu, Shixuan
Lu, Dunjie
Luo, Ruilin
Lv, Chenxu
Men, Rui
Meng, Lingchen
Ren, Xuancheng
Ren, Xingzhang
Song, Sibo
Sun, Yuchong
Tang, Jun
Tu, Jianhong
Wan, Jianqiang
Wang, Peng
Wang, Pengfei
Wang, Qiuyue
Wang, Yuxuan
Xie, Tianbao
Xu, Yiheng
Xu, Haiyang
Xu, Jin
Yang, Zhibo
Yang, Mingkun
Yang, Jianxin
Yang, An
Yu, Bowen
Zhang, Fei
Zhang, Hang
Zhang, Xi
Zheng, Bo
Zhong, Humen
Zhou, Jingren
Zhou, Fan
Zhou, Jing
Zhu, Yuanzhi
Zhu, Ke
author_facet Bai, Shuai
Cai, Yuxuan
Chen, Ruizhe
Chen, Keqin
Chen, Xionghui
Cheng, Zesen
Deng, Lianghao
Ding, Wei
Gao, Chang
Ge, Chunjiang
Ge, Wenbin
Guo, Zhifang
Huang, Qidong
Huang, Jie
Huang, Fei
Hui, Binyuan
Jiang, Shutong
Li, Zhaohai
Li, Mingsheng
Li, Mei
Li, Kaixin
Lin, Zicheng
Lin, Junyang
Liu, Xuejing
Liu, Jiawei
Liu, Chenglong
Liu, Yang
Liu, Dayiheng
Liu, Shixuan
Lu, Dunjie
Luo, Ruilin
Lv, Chenxu
Men, Rui
Meng, Lingchen
Ren, Xuancheng
Ren, Xingzhang
Song, Sibo
Sun, Yuchong
Tang, Jun
Tu, Jianhong
Wan, Jianqiang
Wang, Peng
Wang, Pengfei
Wang, Qiuyue
Wang, Yuxuan
Xie, Tianbao
Xu, Yiheng
Xu, Haiyang
Xu, Jin
Yang, Zhibo
Yang, Mingkun
Yang, Jianxin
Yang, An
Yu, Bowen
Zhang, Fei
Zhang, Hang
Zhang, Xi
Zheng, Bo
Zhong, Humen
Zhou, Jingren
Zhou, Fan
Zhou, Jing
Zhu, Yuanzhi
Zhu, Ke
contents We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21631
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Qwen3-VL Technical Report
Bai, Shuai
Cai, Yuxuan
Chen, Ruizhe
Chen, Keqin
Chen, Xionghui
Cheng, Zesen
Deng, Lianghao
Ding, Wei
Gao, Chang
Ge, Chunjiang
Ge, Wenbin
Guo, Zhifang
Huang, Qidong
Huang, Jie
Huang, Fei
Hui, Binyuan
Jiang, Shutong
Li, Zhaohai
Li, Mingsheng
Li, Mei
Li, Kaixin
Lin, Zicheng
Lin, Junyang
Liu, Xuejing
Liu, Jiawei
Liu, Chenglong
Liu, Yang
Liu, Dayiheng
Liu, Shixuan
Lu, Dunjie
Luo, Ruilin
Lv, Chenxu
Men, Rui
Meng, Lingchen
Ren, Xuancheng
Ren, Xingzhang
Song, Sibo
Sun, Yuchong
Tang, Jun
Tu, Jianhong
Wan, Jianqiang
Wang, Peng
Wang, Pengfei
Wang, Qiuyue
Wang, Yuxuan
Xie, Tianbao
Xu, Yiheng
Xu, Haiyang
Xu, Jin
Yang, Zhibo
Yang, Mingkun
Yang, Jianxin
Yang, An
Yu, Bowen
Zhang, Fei
Zhang, Hang
Zhang, Xi
Zheng, Bo
Zhong, Humen
Zhou, Jingren
Zhou, Fan
Zhou, Jing
Zhu, Yuanzhi
Zhu, Ke
Computer Vision and Pattern Recognition
Artificial Intelligence
We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.
title Qwen3-VL Technical Report
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.21631