UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916746856759296 |
|---|---|
| author | Tian, Rui Gao, Mingfei Xu, Mingze Hu, Jiaming Lu, Jiasen Wu, Zuxuan Yang, Yinfei Dehghan, Afshin |
| author_facet | Tian, Rui Gao, Mingfei Xu, Mingze Hu, Jiaming Lu, Jiasen Wu, Zuxuan Yang, Yinfei Dehghan, Afshin |
| contents | We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14682 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Tian, Rui Gao, Mingfei Xu, Mingze Hu, Jiaming Lu, Jiasen Wu, Zuxuan Yang, Yinfei Dehghan, Afshin Computer Vision and Pattern Recognition We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research. |
| title | UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.14682 |