LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: An, Xiang, Xie, Yin, Yang, Kaicheng, Zhang, Wenkang, Zhao, Xiuwei, Cheng, Zheng, Wang, Yirui, Xu, Songcen, Chen, Changrui, Zhu, Didi, Wu, Chunsheng, Tan, Huajie, Li, Chunyuan, Yang, Jing, Yu, Jie, Wang, Xiyao, Qin, Bin, Wang, Yumeng, Yan, Zizhen, Feng, Ziyong, Liu, Ziwei, Li, Bo, Deng, Jiankang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911317873393664
author An, Xiang
Xie, Yin
Yang, Kaicheng
Zhang, Wenkang
Zhao, Xiuwei
Cheng, Zheng
Wang, Yirui
Xu, Songcen
Chen, Changrui
Zhu, Didi
Wu, Chunsheng
Tan, Huajie
Li, Chunyuan
Yang, Jing
Yu, Jie
Wang, Xiyao
Qin, Bin
Wang, Yumeng
Yan, Zizhen
Feng, Ziyong
Liu, Ziwei
Li, Bo
Deng, Jiankang
author_facet An, Xiang
Xie, Yin
Yang, Kaicheng
Zhang, Wenkang
Zhao, Xiuwei
Cheng, Zheng
Wang, Yirui
Xu, Songcen
Chen, Changrui
Zhu, Didi
Wu, Chunsheng
Tan, Huajie
Li, Chunyuan
Yang, Jing
Yu, Jie
Wang, Xiyao
Qin, Bin
Wang, Yumeng
Yan, Zizhen
Feng, Ziyong
Liu, Ziwei
Li, Bo
Deng, Jiankang
contents We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, LLaVA-OneVision-1.5 provides an open, efficient, and reproducible framework for building high-quality vision-language models entirely from scratch. The LLaVA-OneVision-1.5 release comprises three primary components: (1) Large-Scale Curated Datasets: We construct an 85M concept-balanced pretraining dataset LLaVA-OneVision-1.5-Mid-Traning and a meticulously curated 22M instruction dataset LLaVA-OneVision-1.5-Instruct. (2) Efficient Training Framework: We develop a complete end-to-end efficient training framework leveraging an offline parallel data packing strategy to facilitate the training of LLaVA-OneVision-1.5 within a $16,000 budget. (3) State-of-the-art Performance: Experimental results demonstrate that LLaVA-OneVision-1.5 yields exceptionally competitive performance across a broad range of downstream tasks. Specifically, LLaVA-OneVision-1.5-8B outperforms Qwen2.5-VL-7B on 18 of 27 benchmarks, and LLaVA-OneVision-1.5-4B surpasses Qwen2.5-VL-3B on all 27 benchmarks. (4) RL-based Post-training: We unlock the model's latent potential through a lightweight RL stage, effectively eliciting robust chain-of-thought reasoning to significantly boost performance on complex multimodal reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
An, Xiang
Xie, Yin
Yang, Kaicheng
Zhang, Wenkang
Zhao, Xiuwei
Cheng, Zheng
Wang, Yirui
Xu, Songcen
Chen, Changrui
Zhu, Didi
Wu, Chunsheng
Tan, Huajie
Li, Chunyuan
Yang, Jing
Yu, Jie
Wang, Xiyao
Qin, Bin
Wang, Yumeng
Yan, Zizhen
Feng, Ziyong
Liu, Ziwei
Li, Bo
Deng, Jiankang
Computer Vision and Pattern Recognition
We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, LLaVA-OneVision-1.5 provides an open, efficient, and reproducible framework for building high-quality vision-language models entirely from scratch. The LLaVA-OneVision-1.5 release comprises three primary components: (1) Large-Scale Curated Datasets: We construct an 85M concept-balanced pretraining dataset LLaVA-OneVision-1.5-Mid-Traning and a meticulously curated 22M instruction dataset LLaVA-OneVision-1.5-Instruct. (2) Efficient Training Framework: We develop a complete end-to-end efficient training framework leveraging an offline parallel data packing strategy to facilitate the training of LLaVA-OneVision-1.5 within a $16,000 budget. (3) State-of-the-art Performance: Experimental results demonstrate that LLaVA-OneVision-1.5 yields exceptionally competitive performance across a broad range of downstream tasks. Specifically, LLaVA-OneVision-1.5-8B outperforms Qwen2.5-VL-7B on 18 of 27 benchmarks, and LLaVA-OneVision-1.5-4B surpasses Qwen2.5-VL-3B on all 27 benchmarks. (4) RL-based Post-training: We unlock the model's latent potential through a lightweight RL stage, effectively eliciting robust chain-of-thought reasoning to significantly boost performance on complex multimodal reasoning tasks.
title LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.23661