LatentUMM: Dual Latent Alignment for Unified Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Yinyi, Wang, Wenwen, Bai, Hayes, Savvides, Marios, Wang, Jindong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Corrected Image Generation with Explainable Latent Rewards
by: Luo, Yinyi, et al.
Published: (2026)
by: Luo, Yinyi, et al.
Published: (2026)
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
by: Luo, Yinyi, et al.
Published: (2026)
by: Luo, Yinyi, et al.
Published: (2026)
Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
Image Tokenizer Needs Post-Training
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
ChatUMM: Robust Context Tracking for Conversational Interleaved Generation
by: Dai, Wenxun, et al.
Published: (2026)
by: Dai, Wenxun, et al.
Published: (2026)
Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation
by: Maduabuchi, Chika, et al.
Published: (2025)
by: Maduabuchi, Chika, et al.
Published: (2025)
LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning
by: Xu, Haiying, et al.
Published: (2026)
by: Xu, Haiying, et al.
Published: (2026)
IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
by: Zhai, Yuanhao, et al.
Published: (2024)
by: Zhai, Yuanhao, et al.
Published: (2024)
Latent Harmony: Synergistic Unified UHD Image Restoration via Latent Space Regularization and Controllable Refinement
by: Liu, Yidi, et al.
Published: (2025)
by: Liu, Yidi, et al.
Published: (2025)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
by: Jiang, Houcheng, et al.
Published: (2026)
by: Jiang, Houcheng, et al.
Published: (2026)
Conv-Adapter: Exploring Parameter Efficient Transfer Learning for ConvNets
by: Chen, Hao, et al.
Published: (2022)
by: Chen, Hao, et al.
Published: (2022)
An Embarrassingly Simple Baseline for Imbalanced Semi-Supervised Learning
by: Chen, Hao, et al.
Published: (2022)
by: Chen, Hao, et al.
Published: (2022)
InfoSculpt: Sculpting the Latent Space for Generalized Category Discovery
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Boosting Latent Diffusion Models via Disentangled Representation Alignment
by: Page, John, et al.
Published: (2026)
by: Page, John, et al.
Published: (2026)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
by: Li, Hang, et al.
Published: (2023)
by: Li, Hang, et al.
Published: (2023)
A CLIP-based Uncertainty Modal Modeling (UMM) Framework for Pedestrian Re-Identification in Autonomous Driving
by: Li, Jialin, et al.
Published: (2025)
by: Li, Jialin, et al.
Published: (2025)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
by: He, Chenwei, et al.
Published: (2026)
by: He, Chenwei, et al.
Published: (2026)
Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization
by: Li, Haosen, et al.
Published: (2026)
by: Li, Haosen, et al.
Published: (2026)
Multimodal Latent Language Modeling with Next-Token Diffusion
by: Sun, Yutao, et al.
Published: (2024)
by: Sun, Yutao, et al.
Published: (2024)
StableGuard: Towards Unified Copyright Protection and Tamper Localization in Latent Diffusion Models
by: Yang, Haoxin, et al.
Published: (2025)
by: Yang, Haoxin, et al.
Published: (2025)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
by: Chen, Fangyi, et al.
Published: (2024)
by: Chen, Fangyi, et al.
Published: (2024)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment
by: Guan, Tongfan, et al.
Published: (2025)
by: Guan, Tongfan, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
Inference-time Physics Alignment of Video Generative Models with Latent World Models
by: Yuan, Jianhao, et al.
Published: (2026)
by: Yuan, Jianhao, et al.
Published: (2026)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Multimodal-Conditioned Latent Diffusion Models for Fashion Image Editing
by: Baldrati, Alberto, et al.
Published: (2024)
by: Baldrati, Alberto, et al.
Published: (2024)
Dual-Latent Collaborative Decoding for Fidelity-Perception Balanced Image Compression
by: Mao, Qi, et al.
Published: (2026)
by: Mao, Qi, et al.
Published: (2026)
Latent Feature and Attention Dual Erasure Attack against Multi-View Diffusion Models for 3D Assets Protection
by: Sun, Jingwei, et al.
Published: (2024)
by: Sun, Jingwei, et al.
Published: (2024)
Rigel3D: Rig-aware Latents for Animation-Ready 3D Asset Generation
by: Chatzis, Nikitas, et al.
Published: (2026)
by: Chatzis, Nikitas, et al.
Published: (2026)
UniFL: Improve Latent Diffusion Model via Unified Feedback Learning
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
by: Su, Zhaolong, et al.
Published: (2025)
by: Su, Zhaolong, et al.
Published: (2025)
Motus: A Unified Latent Action World Model
by: Bi, Hongzhe, et al.
Published: (2025)
by: Bi, Hongzhe, et al.
Published: (2025)
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
by: Wang, Chenfeng, et al.
Published: (2026)
by: Wang, Chenfeng, et al.
Published: (2026)
One Latent Space to Rule All Degradations: Unifying Restoration Knowledge for Image Fusion
by: Ma, Haolong, et al.
Published: (2025)
by: Ma, Haolong, et al.
Published: (2025)
HDR Video Generation via Latent Alignment with Logarithmic Encoding
by: Korem, Naomi Ken, et al.
Published: (2026)
by: Korem, Naomi Ken, et al.
Published: (2026)
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
by: Xia, Tianze, et al.
Published: (2025)
by: Xia, Tianze, et al.
Published: (2025)
Similar Items
-
Self-Corrected Image Generation with Explainable Latent Rewards
by: Luo, Yinyi, et al.
Published: (2026) -
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
by: Luo, Yinyi, et al.
Published: (2026) -
Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
by: Qiu, Kai, et al.
Published: (2025) -
Latent Denoising Improves Visual Alignment in Large Multimodal Models
by: Parikh, Dhruv, et al.
Published: (2026) -
Image Tokenizer Needs Post-Training
by: Qiu, Kai, et al.
Published: (2025)