LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911317873393664 |
|---|---|
| author | An, Xiang Xie, Yin Yang, Kaicheng Zhang, Wenkang Zhao, Xiuwei Cheng, Zheng Wang, Yirui Xu, Songcen Chen, Changrui Zhu, Didi Wu, Chunsheng Tan, Huajie Li, Chunyuan Yang, Jing Yu, Jie Wang, Xiyao Qin, Bin Wang, Yumeng Yan, Zizhen Feng, Ziyong Liu, Ziwei Li, Bo Deng, Jiankang |
| author_facet | An, Xiang Xie, Yin Yang, Kaicheng Zhang, Wenkang Zhao, Xiuwei Cheng, Zheng Wang, Yirui Xu, Songcen Chen, Changrui Zhu, Didi Wu, Chunsheng Tan, Huajie Li, Chunyuan Yang, Jing Yu, Jie Wang, Xiyao Qin, Bin Wang, Yumeng Yan, Zizhen Feng, Ziyong Liu, Ziwei Li, Bo Deng, Jiankang |
| contents | We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, LLaVA-OneVision-1.5 provides an open, efficient, and reproducible framework for building high-quality vision-language models entirely from scratch. The LLaVA-OneVision-1.5 release comprises three primary components: (1) Large-Scale Curated Datasets: We construct an 85M concept-balanced pretraining dataset LLaVA-OneVision-1.5-Mid-Traning and a meticulously curated 22M instruction dataset LLaVA-OneVision-1.5-Instruct. (2) Efficient Training Framework: We develop a complete end-to-end efficient training framework leveraging an offline parallel data packing strategy to facilitate the training of LLaVA-OneVision-1.5 within a $16,000 budget. (3) State-of-the-art Performance: Experimental results demonstrate that LLaVA-OneVision-1.5 yields exceptionally competitive performance across a broad range of downstream tasks. Specifically, LLaVA-OneVision-1.5-8B outperforms Qwen2.5-VL-7B on 18 of 27 benchmarks, and LLaVA-OneVision-1.5-4B surpasses Qwen2.5-VL-3B on all 27 benchmarks. (4) RL-based Post-training: We unlock the model's latent potential through a lightweight RL stage, effectively eliciting robust chain-of-thought reasoning to significantly boost performance on complex multimodal reasoning tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_23661 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training An, Xiang Xie, Yin Yang, Kaicheng Zhang, Wenkang Zhao, Xiuwei Cheng, Zheng Wang, Yirui Xu, Songcen Chen, Changrui Zhu, Didi Wu, Chunsheng Tan, Huajie Li, Chunyuan Yang, Jing Yu, Jie Wang, Xiyao Qin, Bin Wang, Yumeng Yan, Zizhen Feng, Ziyong Liu, Ziwei Li, Bo Deng, Jiankang Computer Vision and Pattern Recognition We present LLaVA-OneVision-1.5, a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. Different from the existing works, LLaVA-OneVision-1.5 provides an open, efficient, and reproducible framework for building high-quality vision-language models entirely from scratch. The LLaVA-OneVision-1.5 release comprises three primary components: (1) Large-Scale Curated Datasets: We construct an 85M concept-balanced pretraining dataset LLaVA-OneVision-1.5-Mid-Traning and a meticulously curated 22M instruction dataset LLaVA-OneVision-1.5-Instruct. (2) Efficient Training Framework: We develop a complete end-to-end efficient training framework leveraging an offline parallel data packing strategy to facilitate the training of LLaVA-OneVision-1.5 within a $16,000 budget. (3) State-of-the-art Performance: Experimental results demonstrate that LLaVA-OneVision-1.5 yields exceptionally competitive performance across a broad range of downstream tasks. Specifically, LLaVA-OneVision-1.5-8B outperforms Qwen2.5-VL-7B on 18 of 27 benchmarks, and LLaVA-OneVision-1.5-4B surpasses Qwen2.5-VL-3B on all 27 benchmarks. (4) RL-based Post-training: We unlock the model's latent potential through a lightweight RL stage, effectively eliciting robust chain-of-thought reasoning to significantly boost performance on complex multimodal reasoning tasks. |
| title | LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.23661 |