MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mi, Yapeng, Zhao, Yanpeng, Li, Hengli, Li, Chenxi, Wu, Huimin, Ma, Xiaojian, Zhu, Song-Chun, Wu, Ying Nian, Li, Qing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911436009111552
author Mi, Yapeng
Zhao, Yanpeng
Li, Hengli
Li, Chenxi
Wu, Huimin
Ma, Xiaojian
Zhu, Song-Chun
Wu, Ying Nian
Li, Qing
author_facet Mi, Yapeng
Zhao, Yanpeng
Li, Hengli
Li, Chenxi
Wu, Huimin
Ma, Xiaojian
Zhu, Song-Chun
Wu, Ying Nian
Li, Qing
contents Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image generation either restrict reasoning to a single modality (image or text) or rely on high-quality reasoning data for fine-tuning. To tackle these limitations, we propose MILR, a test-time method that jointly reasons over image and text in a unified latent vector space. Reasoning in MILR is performed by searching through vector representations of discrete image and text tokens. Practically, this is implemented via the policy gradient method, guided by an image quality critic. We instantiate MILR within the unified multimodal understanding and generation (MUG) framework that natively supports language reasoning before image synthesis and thus facilitates cross-modal reasoning. The intermediate model outputs, which are to be optimized, serve as the unified latent space, enabling MILR to operate entirely at test time. We evaluate MILR on GenEval, T2I-CompBench, and WISE, achieving state-of-the-art results on all benchmarks. Notably, on knowledge-intensive WISE, MILR attains an overall score of 0.63, improving over the baseline by 80%. Our further analysis indicates that joint reasoning in the unified latent space is the key to its strong performance. Moreover, our qualitative studies reveal MILR's non-trivial ability in temporal and cultural reasoning, highlighting the efficacy of our reasoning method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22761
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
Mi, Yapeng
Zhao, Yanpeng
Li, Hengli
Li, Chenxi
Wu, Huimin
Ma, Xiaojian
Zhu, Song-Chun
Wu, Ying Nian
Li, Qing
Computer Vision and Pattern Recognition
Artificial Intelligence
Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image generation either restrict reasoning to a single modality (image or text) or rely on high-quality reasoning data for fine-tuning. To tackle these limitations, we propose MILR, a test-time method that jointly reasons over image and text in a unified latent vector space. Reasoning in MILR is performed by searching through vector representations of discrete image and text tokens. Practically, this is implemented via the policy gradient method, guided by an image quality critic. We instantiate MILR within the unified multimodal understanding and generation (MUG) framework that natively supports language reasoning before image synthesis and thus facilitates cross-modal reasoning. The intermediate model outputs, which are to be optimized, serve as the unified latent space, enabling MILR to operate entirely at test time. We evaluate MILR on GenEval, T2I-CompBench, and WISE, achieving state-of-the-art results on all benchmarks. Notably, on knowledge-intensive WISE, MILR attains an overall score of 0.63, improving over the baseline by 80%. Our further analysis indicates that joint reasoning in the unified latent space is the key to its strong performance. Moreover, our qualitative studies reveal MILR's non-trivial ability in temporal and cultural reasoning, highlighting the efficacy of our reasoning method.
title MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.22761