_version_ 1866911124392247296
author Wang, Weiyun
Gao, Zhangwei
Gu, Lixin
Pu, Hengjun
Cui, Long
Wei, Xingguang
Liu, Zhaoyang
Jing, Linglin
Ye, Shenglong
Shao, Jie
Wang, Zhaokai
Chen, Zhe
Zhang, Hongjie
Yang, Ganlin
Wang, Haomin
Wei, Qi
Yin, Jinhui
Li, Wenhao
Cui, Erfei
Chen, Guanzhou
Ding, Zichen
Tian, Changyao
Wu, Zhenyu
Xie, Jingjing
Li, Zehao
Yang, Bowen
Duan, Yuchen
Wang, Xuehui
Hou, Zhi
Hao, Haoran
Zhang, Tianyi
Li, Songze
Zhao, Xiangyu
Duan, Haodong
Deng, Nianchen
Fu, Bin
He, Yinan
Wang, Yi
He, Conghui
Shi, Botian
He, Junjun
Xiong, Yingtong
Lv, Han
Wu, Lijun
Shao, Wenqi
Zhang, Kaipeng
Deng, Huipeng
Qi, Biqing
Ge, Jiaye
Guo, Qipeng
Zhang, Wenwei
Zhang, Songyang
Cao, Maosong
Lin, Junyao
Tang, Kexian
Gao, Jianfei
Huang, Haian
Gu, Yuzhe
Lyu, Chengqi
Tang, Huanze
Wang, Rui
Lv, Haijun
Ouyang, Wanli
Wang, Limin
Dou, Min
Zhu, Xizhou
Lu, Tong
Lin, Dahua
Dai, Jifeng
Su, Weijie
Zhou, Bowen
Chen, Kai
Qiao, Yu
Wang, Wenhai
Luo, Gen
author_facet Wang, Weiyun
Gao, Zhangwei
Gu, Lixin
Pu, Hengjun
Cui, Long
Wei, Xingguang
Liu, Zhaoyang
Jing, Linglin
Ye, Shenglong
Shao, Jie
Wang, Zhaokai
Chen, Zhe
Zhang, Hongjie
Yang, Ganlin
Wang, Haomin
Wei, Qi
Yin, Jinhui
Li, Wenhao
Cui, Erfei
Chen, Guanzhou
Ding, Zichen
Tian, Changyao
Wu, Zhenyu
Xie, Jingjing
Li, Zehao
Yang, Bowen
Duan, Yuchen
Wang, Xuehui
Hou, Zhi
Hao, Haoran
Zhang, Tianyi
Li, Songze
Zhao, Xiangyu
Duan, Haodong
Deng, Nianchen
Fu, Bin
He, Yinan
Wang, Yi
He, Conghui
Shi, Botian
He, Junjun
Xiong, Yingtong
Lv, Han
Wu, Lijun
Shao, Wenqi
Zhang, Kaipeng
Deng, Huipeng
Qi, Biqing
Ge, Jiaye
Guo, Qipeng
Zhang, Wenwei
Zhang, Songyang
Cao, Maosong
Lin, Junyao
Tang, Kexian
Gao, Jianfei
Huang, Haian
Gu, Yuzhe
Lyu, Chengqi
Tang, Huanze
Wang, Rui
Lv, Haijun
Ouyang, Wanli
Wang, Limin
Dou, Min
Zhu, Xizhou
Lu, Tong
Lin, Dahua
Dai, Jifeng
Su, Weijie
Zhou, Bowen
Chen, Kai
Qiao, Yu
Wang, Wenhai
Luo, Gen
contents We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18265
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Wang, Weiyun
Gao, Zhangwei
Gu, Lixin
Pu, Hengjun
Cui, Long
Wei, Xingguang
Liu, Zhaoyang
Jing, Linglin
Ye, Shenglong
Shao, Jie
Wang, Zhaokai
Chen, Zhe
Zhang, Hongjie
Yang, Ganlin
Wang, Haomin
Wei, Qi
Yin, Jinhui
Li, Wenhao
Cui, Erfei
Chen, Guanzhou
Ding, Zichen
Tian, Changyao
Wu, Zhenyu
Xie, Jingjing
Li, Zehao
Yang, Bowen
Duan, Yuchen
Wang, Xuehui
Hou, Zhi
Hao, Haoran
Zhang, Tianyi
Li, Songze
Zhao, Xiangyu
Duan, Haodong
Deng, Nianchen
Fu, Bin
He, Yinan
Wang, Yi
He, Conghui
Shi, Botian
He, Junjun
Xiong, Yingtong
Lv, Han
Wu, Lijun
Shao, Wenqi
Zhang, Kaipeng
Deng, Huipeng
Qi, Biqing
Ge, Jiaye
Guo, Qipeng
Zhang, Wenwei
Zhang, Songyang
Cao, Maosong
Lin, Junyao
Tang, Kexian
Gao, Jianfei
Huang, Haian
Gu, Yuzhe
Lyu, Chengqi
Tang, Huanze
Wang, Rui
Lv, Haijun
Ouyang, Wanli
Wang, Limin
Dou, Min
Zhu, Xizhou
Lu, Tong
Lin, Dahua
Dai, Jifeng
Su, Weijie
Zhou, Bowen
Chen, Kai
Qiao, Yu
Wang, Wenhai
Luo, Gen
Computer Vision and Pattern Recognition
We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released.
title InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.18265