LatentUMM: Dual Latent Alignment for Unified Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Yinyi, Wang, Wenwen, Bai, Hayes, Savvides, Marios, Wang, Jindong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914576239427584
author Luo, Yinyi
Wang, Wenwen
Bai, Hayes
Savvides, Marios
Wang, Jindong
author_facet Luo, Yinyi
Wang, Wenwen
Bai, Hayes
Savvides, Marios
Wang, Jindong
contents Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We observe that this issue does not stem from a lack of shared representations, but from the absence of explicit alignment between the transformations that map into and out of the latent space. As a result, generation and re-encoding can follow inconsistent trajectories, leading to semantic drift under modality transitions. In this work, we propose LatentUMM, a framework that constructs an enhanced shared latent space to explicitly align these transformations and improve cross-modal consistency. LatentUMM consists of two stages. First, dual latent alignment enforces consistency at both the modality and capacity levels: cross-modal alignment uses a stronger embedding model to impose structured cross-modal semantics, while dual capacity alignment enforces bidirectional consistency under generation and re-encoding. Second, latent dynamics stabilization improves robustness via stochastic latent rollouts and preference optimization, favoring trajectories that better preserve semantic consistency. Experiments show that LatentUMM consistently improves multimodal consistency across diverse architectures. Code is available at: https://github.com/AIFrontierLab/TorchUMM/tree/main/src/umm/post_training/LatentUMM.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17766
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LatentUMM: Dual Latent Alignment for Unified Multimodal Models
Luo, Yinyi
Wang, Wenwen
Bai, Hayes
Savvides, Marios
Wang, Jindong
Computer Vision and Pattern Recognition
Unified multimodal models (UMMs) achieve strong performance in both understanding and generation by learning a shared latent space, yet they often exhibit functional inconsistency between these two capabilities. We observe that this issue does not stem from a lack of shared representations, but from the absence of explicit alignment between the transformations that map into and out of the latent space. As a result, generation and re-encoding can follow inconsistent trajectories, leading to semantic drift under modality transitions. In this work, we propose LatentUMM, a framework that constructs an enhanced shared latent space to explicitly align these transformations and improve cross-modal consistency. LatentUMM consists of two stages. First, dual latent alignment enforces consistency at both the modality and capacity levels: cross-modal alignment uses a stronger embedding model to impose structured cross-modal semantics, while dual capacity alignment enforces bidirectional consistency under generation and re-encoding. Second, latent dynamics stabilization improves robustness via stochastic latent rollouts and preference optimization, favoring trajectories that better preserve semantic consistency. Experiments show that LatentUMM consistently improves multimodal consistency across diverse architectures. Code is available at: https://github.com/AIFrontierLab/TorchUMM/tree/main/src/umm/post_training/LatentUMM.
title LatentUMM: Dual Latent Alignment for Unified Multimodal Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17766