Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912600031232000 |
|---|---|
| author | Dong, Daxiang Zheng, Mingming Xu, Dong Zhuang, Bairong Zhang, Wenyu Luo, Chunhua Wang, Haoran Zhao, Zijian Li, Jie Li, Yuxuan Zhong, Hanjun Liu, Mengyue Chen, Jieting Li, Shupeng Tian, Lun Feng, Yaping Li, Xin Jiang, Donggang Chen, Yong Xu, Yehua Qin, Duohao Feng, Chen Wang, Dan Zhang, Henghua Ha, Jingjing He, Jinhui Zhai, Yanfeng Zheng, Chengxin Mao, Jiayi Chen, Jiacheng Yao, Ruchang Yuan, Ziye Wu, Jianmin Xie, Guangjun Shen, Dou |
| author_facet | Dong, Daxiang Zheng, Mingming Xu, Dong Zhuang, Bairong Zhang, Wenyu Luo, Chunhua Wang, Haoran Zhao, Zijian Li, Jie Li, Yuxuan Zhong, Hanjun Liu, Mengyue Chen, Jieting Li, Shupeng Tian, Lun Feng, Yaping Li, Xin Jiang, Donggang Chen, Yong Xu, Yehua Qin, Duohao Feng, Chen Wang, Dan Zhang, Henghua Ha, Jingjing He, Jinhui Zhai, Yanfeng Zheng, Chengxin Mao, Jiayi Chen, Jiacheng Yao, Ruchang Yuan, Ziye Wu, Jianmin Xie, Guangjun Shen, Dou |
| contents | We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement techniques. Our approach employs multi-stage progressive training and high-precision data synthesis pipelines, which prove to be critical technologies for enhancing domain-specific capabilities while maintaining strong general performance. Qianfan-VL achieves comparable results to leading open-source models on general benchmarks, with state-of-the-art performance on benchmarks such as CCBench, SEEDBench IMG, ScienceQA, and MMStar. The domain enhancement strategy delivers significant advantages in OCR and document understanding, validated on both public benchmarks (OCRBench 873, DocVQA 94.75%) and in-house evaluations. Notably, Qianfan-VL-8B and 70B variants incorporate long chain-of-thought capabilities, demonstrating superior performance on mathematical reasoning (MathVista 78.6%) and logical inference tasks. All models are trained entirely on Baidu's Kunlun P800 chips, validating the capability of large-scale AI infrastructure to train SOTA-level multimodal models with over 90% scaling efficiency on 5000 chips for a single task. This work establishes an effective methodology for developing domain-enhanced multimodal models suitable for diverse enterprise deployment scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_18189 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Qianfan-VL: Domain-Enhanced Universal Vision-Language Models Dong, Daxiang Zheng, Mingming Xu, Dong Zhuang, Bairong Zhang, Wenyu Luo, Chunhua Wang, Haoran Zhao, Zijian Li, Jie Li, Yuxuan Zhong, Hanjun Liu, Mengyue Chen, Jieting Li, Shupeng Tian, Lun Feng, Yaping Li, Xin Jiang, Donggang Chen, Yong Xu, Yehua Qin, Duohao Feng, Chen Wang, Dan Zhang, Henghua Ha, Jingjing He, Jinhui Zhai, Yanfeng Zheng, Chengxin Mao, Jiayi Chen, Jiacheng Yao, Ruchang Yuan, Ziye Wu, Jianmin Xie, Guangjun Shen, Dou Computer Vision and Pattern Recognition Artificial Intelligence We present Qianfan-VL, a series of multimodal large language models ranging from 3B to 70B parameters, achieving state-of-the-art performance through innovative domain enhancement techniques. Our approach employs multi-stage progressive training and high-precision data synthesis pipelines, which prove to be critical technologies for enhancing domain-specific capabilities while maintaining strong general performance. Qianfan-VL achieves comparable results to leading open-source models on general benchmarks, with state-of-the-art performance on benchmarks such as CCBench, SEEDBench IMG, ScienceQA, and MMStar. The domain enhancement strategy delivers significant advantages in OCR and document understanding, validated on both public benchmarks (OCRBench 873, DocVQA 94.75%) and in-house evaluations. Notably, Qianfan-VL-8B and 70B variants incorporate long chain-of-thought capabilities, demonstrating superior performance on mathematical reasoning (MathVista 78.6%) and logical inference tasks. All models are trained entirely on Baidu's Kunlun P800 chips, validating the capability of large-scale AI infrastructure to train SOTA-level multimodal models with over 90% scaling efficiency on 5000 chips for a single task. This work establishes an effective methodology for developing domain-enhanced multimodal models suitable for diverse enterprise deployment scenarios. |
| title | Qianfan-VL: Domain-Enhanced Universal Vision-Language Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2509.18189 |