UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tian, Rui, Gao, Mingfei, Xu, Mingze, Hu, Jiaming, Lu, Jiasen, Wu, Zuxuan, Yang, Yinfei, Dehghan, Afshin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916746856759296
author Tian, Rui
Gao, Mingfei
Xu, Mingze
Hu, Jiaming
Lu, Jiasen
Wu, Zuxuan
Yang, Yinfei
Dehghan, Afshin
author_facet Tian, Rui
Gao, Mingfei
Xu, Mingze
Hu, Jiaming
Lu, Jiasen
Wu, Zuxuan
Yang, Yinfei
Dehghan, Afshin
contents We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14682
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Tian, Rui
Gao, Mingfei
Xu, Mingze
Hu, Jiaming
Lu, Jiasen
Wu, Zuxuan
Yang, Yinfei
Dehghan, Afshin
Computer Vision and Pattern Recognition
We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training, supervised fine-tuning, and direct preference optimization. More importantly, we propose a new Chain-of-Thought Verification (CoT-V) strategy for test-time scaling, which significantly boosts UniGen's image generation quality using a simple Best-of-N test-time strategy. Specifically, CoT-V enables UniGen to act as both image generator and verifier at test time, assessing the semantic alignment between a text prompt and its generated image in a step-by-step CoT manner. Trained entirely on open-source datasets across all stages, UniGen achieves state-of-the-art performance on a range of image understanding and generation benchmarks, with a final score of 0.78 on GenEval and 85.19 on DPG-Bench. Through extensive ablation studies, our work provides actionable insights and addresses key challenges in the full life cycle of building unified MLLMs, contributing meaningful directions to the future research.
title UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.14682