Qwen3-VL Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917109959753728 |
|---|---|
| author | Bai, Shuai Cai, Yuxuan Chen, Ruizhe Chen, Keqin Chen, Xionghui Cheng, Zesen Deng, Lianghao Ding, Wei Gao, Chang Ge, Chunjiang Ge, Wenbin Guo, Zhifang Huang, Qidong Huang, Jie Huang, Fei Hui, Binyuan Jiang, Shutong Li, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Lin, Zicheng Lin, Junyang Liu, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Lu, Dunjie Luo, Ruilin Lv, Chenxu Men, Rui Meng, Lingchen Ren, Xuancheng Ren, Xingzhang Song, Sibo Sun, Yuchong Tang, Jun Tu, Jianhong Wan, Jianqiang Wang, Peng Wang, Pengfei Wang, Qiuyue Wang, Yuxuan Xie, Tianbao Xu, Yiheng Xu, Haiyang Xu, Jin Yang, Zhibo Yang, Mingkun Yang, Jianxin Yang, An Yu, Bowen Zhang, Fei Zhang, Hang Zhang, Xi Zheng, Bo Zhong, Humen Zhou, Jingren Zhou, Fan Zhou, Jing Zhu, Yuanzhi Zhu, Ke |
| author_facet | Bai, Shuai Cai, Yuxuan Chen, Ruizhe Chen, Keqin Chen, Xionghui Cheng, Zesen Deng, Lianghao Ding, Wei Gao, Chang Ge, Chunjiang Ge, Wenbin Guo, Zhifang Huang, Qidong Huang, Jie Huang, Fei Hui, Binyuan Jiang, Shutong Li, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Lin, Zicheng Lin, Junyang Liu, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Lu, Dunjie Luo, Ruilin Lv, Chenxu Men, Rui Meng, Lingchen Ren, Xuancheng Ren, Xingzhang Song, Sibo Sun, Yuchong Tang, Jun Tu, Jianhong Wan, Jianqiang Wang, Peng Wang, Pengfei Wang, Qiuyue Wang, Yuxuan Xie, Tianbao Xu, Yiheng Xu, Haiyang Xu, Jin Yang, Zhibo Yang, Mingkun Yang, Jianxin Yang, An Yu, Bowen Zhang, Fei Zhang, Hang Zhang, Xi Zheng, Bo Zhong, Humen Zhou, Jingren Zhou, Fan Zhou, Jing Zhu, Yuanzhi Zhu, Ke |
| contents | We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_21631 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Qwen3-VL Technical Report Bai, Shuai Cai, Yuxuan Chen, Ruizhe Chen, Keqin Chen, Xionghui Cheng, Zesen Deng, Lianghao Ding, Wei Gao, Chang Ge, Chunjiang Ge, Wenbin Guo, Zhifang Huang, Qidong Huang, Jie Huang, Fei Hui, Binyuan Jiang, Shutong Li, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Lin, Zicheng Lin, Junyang Liu, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Lu, Dunjie Luo, Ruilin Lv, Chenxu Men, Rui Meng, Lingchen Ren, Xuancheng Ren, Xingzhang Song, Sibo Sun, Yuchong Tang, Jun Tu, Jianhong Wan, Jianqiang Wang, Peng Wang, Pengfei Wang, Qiuyue Wang, Yuxuan Xie, Tianbao Xu, Yiheng Xu, Haiyang Xu, Jin Yang, Zhibo Yang, Mingkun Yang, Jianxin Yang, An Yu, Bowen Zhang, Fei Zhang, Hang Zhang, Xi Zheng, Bo Zhong, Humen Zhou, Jingren Zhou, Fan Zhou, Jing Zhu, Yuanzhi Zhu, Ke Computer Vision and Pattern Recognition Artificial Intelligence We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows. |
| title | Qwen3-VL Technical Report |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2511.21631 |