Saved in:
Bibliographic Details
Main Authors: Wang, Chenlong, Chen, Yuhang, Hu, Zhihan, Chen, Dongping, Chen, Wenhu, Wiegreffe, Sarah, Zhou, Tianyi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.02140
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910008806998016
author Wang, Chenlong
Chen, Yuhang
Hu, Zhihan
Chen, Dongping
Chen, Wenhu
Wiegreffe, Sarah
Zhou, Tianyi
author_facet Wang, Chenlong
Chen, Yuhang
Hu, Zhihan
Chen, Dongping
Chen, Wenhu
Wiegreffe, Sarah
Zhou, Tianyi
contents Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are genuinely aligned and integrated within a single model remains unclear. To investigate this question, we introduce GapEval, a bidirectional benchmark designed to quantify the gap between understanding and generation capabilities, and quantitatively measure the cognitive coherence of the two "unified" directions. Each question can be answered in both modalities (image and text), enabling a symmetric evaluation of a model's bidirectional inference capability and cross-modal consistency. Experiments reveal a persistent gap between the two directions across a wide range of UMMs with different architectures, suggesting that current models achieve only surface-level unification rather than deep cognitive convergence of the two. To further explore the underlying mechanism, we conduct an empirical study from the perspective of knowledge manipulation to illustrate the underlying limitations. Our findings indicate that knowledge within UMMs often remains disjoint. The capability emergence and knowledge across modalities are unsynchronized, paving the way for further exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02140
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
Wang, Chenlong
Chen, Yuhang
Hu, Zhihan
Chen, Dongping
Chen, Wenhu
Wiegreffe, Sarah
Zhou, Tianyi
Computation and Language
Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are genuinely aligned and integrated within a single model remains unclear. To investigate this question, we introduce GapEval, a bidirectional benchmark designed to quantify the gap between understanding and generation capabilities, and quantitatively measure the cognitive coherence of the two "unified" directions. Each question can be answered in both modalities (image and text), enabling a symmetric evaluation of a model's bidirectional inference capability and cross-modal consistency. Experiments reveal a persistent gap between the two directions across a wide range of UMMs with different architectures, suggesting that current models achieve only surface-level unification rather than deep cognitive convergence of the two. To further explore the underlying mechanism, we conduct an empirical study from the perspective of knowledge manipulation to illustrate the underlying limitations. Our findings indicate that knowledge within UMMs often remains disjoint. The capability emergence and knowledge across modalities are unsynchronized, paving the way for further exploration.
title Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
topic Computation and Language
url https://arxiv.org/abs/2602.02140