OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Size, Wu, Zhonghua, Gong, Zerui, Tao, Qingyi, Jin, Sheng, Li, Qinyue, Li, Wei, Loy, Chen Change
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913869535903744
author Wu, Size
Wu, Zhonghua
Gong, Zerui
Tao, Qingyi
Jin, Sheng
Li, Qinyue
Li, Wei
Loy, Chen Change
author_facet Wu, Size
Wu, Zhonghua
Gong, Zerui
Tao, Qingyi
Jin, Sheng
Li, Qinyue
Li, Wei
Loy, Chen Change
contents In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training strategy that minimizes the training complexity and overhead by bridging the off-the-shelf multimodal large language models (LLMs) and diffusion models through a set of learnable queries and a light-weight transformer-based connector. With a minimalist choice of architecture, we demonstrate that OpenUni can: 1) generate high-quality and instruction-aligned images, and 2) achieve exceptional performance on standard benchmarks such as GenEval, DPG- Bench, and WISE, with only 1.1B and 3.1B activated parameters. To support open research and community advancement, we release all model weights, training code, and our curated training datasets (including 23M image-text pairs) at https://github.com/wusize/OpenUni.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Wu, Size
Wu, Zhonghua
Gong, Zerui
Tao, Qingyi
Jin, Sheng
Li, Qinyue
Li, Wei
Loy, Chen Change
Computer Vision and Pattern Recognition
In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training strategy that minimizes the training complexity and overhead by bridging the off-the-shelf multimodal large language models (LLMs) and diffusion models through a set of learnable queries and a light-weight transformer-based connector. With a minimalist choice of architecture, we demonstrate that OpenUni can: 1) generate high-quality and instruction-aligned images, and 2) achieve exceptional performance on standard benchmarks such as GenEval, DPG- Bench, and WISE, with only 1.1B and 3.1B activated parameters. To support open research and community advancement, we release all model weights, training code, and our curated training datasets (including 23M image-text pairs) at https://github.com/wusize/OpenUni.
title OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.23661