InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911124392247296 |
|---|---|
| author | Wang, Weiyun Gao, Zhangwei Gu, Lixin Pu, Hengjun Cui, Long Wei, Xingguang Liu, Zhaoyang Jing, Linglin Ye, Shenglong Shao, Jie Wang, Zhaokai Chen, Zhe Zhang, Hongjie Yang, Ganlin Wang, Haomin Wei, Qi Yin, Jinhui Li, Wenhao Cui, Erfei Chen, Guanzhou Ding, Zichen Tian, Changyao Wu, Zhenyu Xie, Jingjing Li, Zehao Yang, Bowen Duan, Yuchen Wang, Xuehui Hou, Zhi Hao, Haoran Zhang, Tianyi Li, Songze Zhao, Xiangyu Duan, Haodong Deng, Nianchen Fu, Bin He, Yinan Wang, Yi He, Conghui Shi, Botian He, Junjun Xiong, Yingtong Lv, Han Wu, Lijun Shao, Wenqi Zhang, Kaipeng Deng, Huipeng Qi, Biqing Ge, Jiaye Guo, Qipeng Zhang, Wenwei Zhang, Songyang Cao, Maosong Lin, Junyao Tang, Kexian Gao, Jianfei Huang, Haian Gu, Yuzhe Lyu, Chengqi Tang, Huanze Wang, Rui Lv, Haijun Ouyang, Wanli Wang, Limin Dou, Min Zhu, Xizhou Lu, Tong Lin, Dahua Dai, Jifeng Su, Weijie Zhou, Bowen Chen, Kai Qiao, Yu Wang, Wenhai Luo, Gen |
| author_facet | Wang, Weiyun Gao, Zhangwei Gu, Lixin Pu, Hengjun Cui, Long Wei, Xingguang Liu, Zhaoyang Jing, Linglin Ye, Shenglong Shao, Jie Wang, Zhaokai Chen, Zhe Zhang, Hongjie Yang, Ganlin Wang, Haomin Wei, Qi Yin, Jinhui Li, Wenhao Cui, Erfei Chen, Guanzhou Ding, Zichen Tian, Changyao Wu, Zhenyu Xie, Jingjing Li, Zehao Yang, Bowen Duan, Yuchen Wang, Xuehui Hou, Zhi Hao, Haoran Zhang, Tianyi Li, Songze Zhao, Xiangyu Duan, Haodong Deng, Nianchen Fu, Bin He, Yinan Wang, Yi He, Conghui Shi, Botian He, Junjun Xiong, Yingtong Lv, Han Wu, Lijun Shao, Wenqi Zhang, Kaipeng Deng, Huipeng Qi, Biqing Ge, Jiaye Guo, Qipeng Zhang, Wenwei Zhang, Songyang Cao, Maosong Lin, Junyao Tang, Kexian Gao, Jianfei Huang, Haian Gu, Yuzhe Lyu, Chengqi Tang, Huanze Wang, Rui Lv, Haijun Ouyang, Wanli Wang, Limin Dou, Min Zhu, Xizhou Lu, Tong Lin, Dahua Dai, Jifeng Su, Weijie Zhou, Bowen Chen, Kai Qiao, Yu Wang, Wenhai Luo, Gen |
| contents | We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_18265 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Wang, Weiyun Gao, Zhangwei Gu, Lixin Pu, Hengjun Cui, Long Wei, Xingguang Liu, Zhaoyang Jing, Linglin Ye, Shenglong Shao, Jie Wang, Zhaokai Chen, Zhe Zhang, Hongjie Yang, Ganlin Wang, Haomin Wei, Qi Yin, Jinhui Li, Wenhao Cui, Erfei Chen, Guanzhou Ding, Zichen Tian, Changyao Wu, Zhenyu Xie, Jingjing Li, Zehao Yang, Bowen Duan, Yuchen Wang, Xuehui Hou, Zhi Hao, Haoran Zhang, Tianyi Li, Songze Zhao, Xiangyu Duan, Haodong Deng, Nianchen Fu, Bin He, Yinan Wang, Yi He, Conghui Shi, Botian He, Junjun Xiong, Yingtong Lv, Han Wu, Lijun Shao, Wenqi Zhang, Kaipeng Deng, Huipeng Qi, Biqing Ge, Jiaye Guo, Qipeng Zhang, Wenwei Zhang, Songyang Cao, Maosong Lin, Junyao Tang, Kexian Gao, Jianfei Huang, Haian Gu, Yuzhe Lyu, Chengqi Tang, Huanze Wang, Rui Lv, Haijun Ouyang, Wanli Wang, Limin Dou, Min Zhu, Xizhou Lu, Tong Lin, Dahua Dai, Jifeng Su, Weijie Zhou, Bowen Chen, Kai Qiao, Yu Wang, Wenhai Luo, Gen Computer Vision and Pattern Recognition We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL series. A key innovation is the Cascade Reinforcement Learning (Cascade RL) framework, which enhances reasoning through a two-stage process: offline RL for stable convergence and online RL for refined alignment. This coarse-to-fine training strategy leads to substantial improvements on downstream reasoning tasks, e.g., MMMU and MathVista. To optimize efficiency, we propose a Visual Resolution Router (ViR) that dynamically adjusts the resolution of visual tokens without compromising performance. Coupled with ViR, our Decoupled Vision-Language Deployment (DvD) strategy separates the vision encoder and language model across different GPUs, effectively balancing computational load. These contributions collectively enable InternVL3.5 to achieve up to a +16.0\% gain in overall reasoning performance and a 4.05$\times$ inference speedup compared to its predecessor, i.e., InternVL3. In addition, InternVL3.5 supports novel capabilities such as GUI interaction and embodied agency. Notably, our largest model, i.e., InternVL3.5-241B-A28B, attains state-of-the-art results among open-source MLLMs across general multimodal, reasoning, text, and agentic tasks -- narrowing the performance gap with leading commercial models like GPT-5. All models and code are publicly released. |
| title | InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.18265 |