MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Tianyu, Wang, Zefan, Wang, Chongyi, Huang, Fuwei, Ma, Wenshuo, He, Zhihui, Cai, Tianchi, Chen, Weize, Huang, Yuxiang, Zhao, Yuanqian, Xu, Bokai, Cui, Junbo, Xu, Yingjing, Ruan, Liqing, Zhang, Luoyuan, Liu, Hanyu, Tang, Jingkun, Liu, Hongyuan, Guo, Qining, Hu, Wenhao, He, Bingxiang, Zhou, Jie, Cai, Jie, Qi, Ji, Guo, Zonghao, Chen, Chi, Zeng, Guoyang, Li, Yuxuan, Cui, Ganqu, Ding, Ning, Han, Xu, Yao, Yuan, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914051307601920
author Yu, Tianyu
Wang, Zefan
Wang, Chongyi
Huang, Fuwei
Ma, Wenshuo
He, Zhihui
Cai, Tianchi
Chen, Weize
Huang, Yuxiang
Zhao, Yuanqian
Xu, Bokai
Cui, Junbo
Xu, Yingjing
Ruan, Liqing
Zhang, Luoyuan
Liu, Hanyu
Tang, Jingkun
Liu, Hongyuan
Guo, Qining
Hu, Wenhao
He, Bingxiang
Zhou, Jie
Cai, Jie
Qi, Ji
Guo, Zonghao
Chen, Chi
Zeng, Guoyang
Li, Yuxuan
Cui, Ganqu
Ding, Ning
Han, Xu
Yao, Yuan
Liu, Zhiyuan
Sun, Maosong
author_facet Yu, Tianyu
Wang, Zefan
Wang, Chongyi
Huang, Fuwei
Ma, Wenshuo
He, Zhihui
Cai, Tianchi
Chen, Weize
Huang, Yuxiang
Zhao, Yuanqian
Xu, Bokai
Cui, Junbo
Xu, Yingjing
Ruan, Liqing
Zhang, Luoyuan
Liu, Hanyu
Tang, Jingkun
Liu, Hongyuan
Guo, Qining
Hu, Wenhao
He, Bingxiang
Zhou, Jie
Cai, Jie
Qi, Ji
Guo, Zonghao
Chen, Chi
Zeng, Guoyang
Li, Yuxuan
Cui, Ganqu
Ding, Ning
Han, Xu
Yao, Yuan
Liu, Zhiyuan
Sun, Maosong
contents Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and scalable. To address the challenges, we present MiniCPM-V 4.5, an 8B parameter model designed for high efficiency and strong performance. We introduce three core improvements in model architecture, data strategy and training method: a unified 3D-Resampler model architecture for highly compact encoding over images and videos, a unified learning paradigm for document knowledge and text recognition without heavy data engineering, and a hybrid reinforcement learning strategy for proficiency in both short and long reasoning modes. Comprehensive experimental results in OpenCompass evaluation show that MiniCPM-V 4.5 surpasses widely used proprietary models such as GPT-4o-latest, and significantly larger open-source models such as Qwen2.5-VL 72B. Notably, the strong performance is achieved with remarkable efficiency. For example, on the widely adopted VideoMME benchmark, MiniCPM-V 4.5 achieves state-of-the-art performance among models under 30B size, using just 46.7\% GPU memory cost and 8.7\% inference time of Qwen2.5-VL 7B.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18154
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Yu, Tianyu
Wang, Zefan
Wang, Chongyi
Huang, Fuwei
Ma, Wenshuo
He, Zhihui
Cai, Tianchi
Chen, Weize
Huang, Yuxiang
Zhao, Yuanqian
Xu, Bokai
Cui, Junbo
Xu, Yingjing
Ruan, Liqing
Zhang, Luoyuan
Liu, Hanyu
Tang, Jingkun
Liu, Hongyuan
Guo, Qining
Hu, Wenhao
He, Bingxiang
Zhou, Jie
Cai, Jie
Qi, Ji
Guo, Zonghao
Chen, Chi
Zeng, Guoyang
Li, Yuxuan
Cui, Ganqu
Ding, Ning
Han, Xu
Yao, Yuan
Liu, Zhiyuan
Sun, Maosong
Machine Learning
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and scalable. To address the challenges, we present MiniCPM-V 4.5, an 8B parameter model designed for high efficiency and strong performance. We introduce three core improvements in model architecture, data strategy and training method: a unified 3D-Resampler model architecture for highly compact encoding over images and videos, a unified learning paradigm for document knowledge and text recognition without heavy data engineering, and a hybrid reinforcement learning strategy for proficiency in both short and long reasoning modes. Comprehensive experimental results in OpenCompass evaluation show that MiniCPM-V 4.5 surpasses widely used proprietary models such as GPT-4o-latest, and significantly larger open-source models such as Qwen2.5-VL 72B. Notably, the strong performance is achieved with remarkable efficiency. For example, on the widely adopted VideoMME benchmark, MiniCPM-V 4.5 achieves state-of-the-art performance among models under 30B size, using just 46.7\% GPU memory cost and 8.7\% inference time of Qwen2.5-VL 7B.
title MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.18154