Mixture of States: Routing Token-Level Dynamics for Multimodal Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Haozhe, Liu, Ding, Zhuge, Mingchen, Zhou, Zijian, Xie, Tian, He, Sen, Yang, Yukang, Liu, Shuming, Cong, Yuren, Guo, Jiadong, Xu, Hongyu, Xu, Ke, Ng, Kam-Woh, Pérez, Juan C., Pérez-Rúa, Juan-Manuel, Xiang, Tao, Liu, Wei, Liu, Shikun, Schmidhuber, Jürgen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914392242651136
author Liu, Haozhe
Liu, Ding
Zhuge, Mingchen
Zhou, Zijian
Xie, Tian
He, Sen
Yang, Yukang
Liu, Shuming
Cong, Yuren
Guo, Jiadong
Xu, Hongyu
Xu, Ke
Ng, Kam-Woh
Pérez, Juan C.
Pérez-Rúa, Juan-Manuel
Xiang, Tao
Liu, Wei
Liu, Shikun
Schmidhuber, Jürgen
author_facet Liu, Haozhe
Liu, Ding
Zhuge, Mingchen
Zhou, Zijian
Xie, Tian
He, Sen
Yang, Yukang
Liu, Shuming
Cong, Yuren
Guo, Jiadong
Xu, Hongyu
Xu, Ke
Ng, Kam-Woh
Pérez, Juan C.
Pérez-Rúa, Juan-Manuel
Xiang, Tao
Liu, Wei
Liu, Shikun
Schmidhuber, Jürgen
contents We introduce MoS (Mixture of States), a novel fusion paradigm for multimodal diffusion models that merges modalities using flexible, state-based interactions. The core of MoS is a learnable, token-wise router that creates denoising timestep- and input-dependent interactions between modalities' hidden states, precisely aligning token-level features with the diffusion trajectory. This router sparsely selects the top-$k$ hidden states and is trained with an $ε$-greedy strategy, efficiently selecting contextual features with minimal learnable parameters and negligible computational overhead. We validate our design with text-to-image generation (MoS-Image) and editing (MoS-Editing), which achieve state-of-the-art results. With only 3B to 5B parameters, our models match or surpass counterparts up to $4\times$ larger. These findings establish MoS as a flexible and compute-efficient paradigm for scaling multimodal diffusion models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12207
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
Liu, Haozhe
Liu, Ding
Zhuge, Mingchen
Zhou, Zijian
Xie, Tian
He, Sen
Yang, Yukang
Liu, Shuming
Cong, Yuren
Guo, Jiadong
Xu, Hongyu
Xu, Ke
Ng, Kam-Woh
Pérez, Juan C.
Pérez-Rúa, Juan-Manuel
Xiang, Tao
Liu, Wei
Liu, Shikun
Schmidhuber, Jürgen
Computer Vision and Pattern Recognition
We introduce MoS (Mixture of States), a novel fusion paradigm for multimodal diffusion models that merges modalities using flexible, state-based interactions. The core of MoS is a learnable, token-wise router that creates denoising timestep- and input-dependent interactions between modalities' hidden states, precisely aligning token-level features with the diffusion trajectory. This router sparsely selects the top-$k$ hidden states and is trained with an $ε$-greedy strategy, efficiently selecting contextual features with minimal learnable parameters and negligible computational overhead. We validate our design with text-to-image generation (MoS-Image) and editing (MoS-Editing), which achieve state-of-the-art results. With only 3B to 5B parameters, our models match or surpass counterparts up to $4\times$ larger. These findings establish MoS as a flexible and compute-efficient paradigm for scaling multimodal diffusion models.
title Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12207