Saved in:
Bibliographic Details
Main Authors: Niu, Yuwei, Jin, Weiyang, Liao, Jiaqi, Feng, Chaoran, Jin, Peng, Lin, Bin, Li, Zongjian, Zhu, Bin, Yu, Weihao, Yuan, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.20561
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912737557217280
author Niu, Yuwei
Jin, Weiyang
Liao, Jiaqi
Feng, Chaoran
Jin, Peng
Lin, Bin
Li, Zongjian
Zhu, Bin
Yu, Weihao
Yuan, Li
author_facet Niu, Yuwei
Jin, Weiyang
Liao, Jiaqi
Feng, Chaoran
Jin, Peng
Lin, Bin
Li, Zongjian
Zhu, Bin
Yu, Weihao
Yuan, Li
contents Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation? To investigate this, we introduce UniSandbox, a decoupled evaluation framework paired with controlled, synthetic datasets to avoid data leakage and enable detailed analysis. Our findings reveal a significant understanding-generation gap, which is mainly reflected in two key dimensions: reasoning generation and knowledge transfer. Specifically, for reasoning generation tasks, we observe that explicit Chain-of-Thought (CoT) in the understanding module effectively bridges the gap, and further demonstrate that a self-training approach can successfully internalize this ability, enabling implicit reasoning during generation. Additionally, for knowledge transfer tasks, we find that CoT assists the generative process by helping retrieve newly learned knowledge, and also discover that query-based architectures inherently exhibit latent CoT-like properties that affect this transfer. UniSandbox provides preliminary insights for designing future unified architectures and training strategies that truly bridge the gap between understanding and generation. Code and data are available at https://github.com/PKU-YuanGroup/UniSandBox
format Preprint
id arxiv_https___arxiv_org_abs_2511_20561
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
Niu, Yuwei
Jin, Weiyang
Liao, Jiaqi
Feng, Chaoran
Jin, Peng
Lin, Bin
Li, Zongjian
Zhu, Bin
Yu, Weihao
Yuan, Li
Computer Vision and Pattern Recognition
Computation and Language
Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation? To investigate this, we introduce UniSandbox, a decoupled evaluation framework paired with controlled, synthetic datasets to avoid data leakage and enable detailed analysis. Our findings reveal a significant understanding-generation gap, which is mainly reflected in two key dimensions: reasoning generation and knowledge transfer. Specifically, for reasoning generation tasks, we observe that explicit Chain-of-Thought (CoT) in the understanding module effectively bridges the gap, and further demonstrate that a self-training approach can successfully internalize this ability, enabling implicit reasoning during generation. Additionally, for knowledge transfer tasks, we find that CoT assists the generative process by helping retrieve newly learned knowledge, and also discover that query-based architectures inherently exhibit latent CoT-like properties that affect this transfer. UniSandbox provides preliminary insights for designing future unified architectures and training strategies that truly bridge the gap between understanding and generation. Code and data are available at https://github.com/PKU-YuanGroup/UniSandBox
title Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2511.20561