Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914392242651136 |
|---|---|
| author | Liu, Haozhe Liu, Ding Zhuge, Mingchen Zhou, Zijian Xie, Tian He, Sen Yang, Yukang Liu, Shuming Cong, Yuren Guo, Jiadong Xu, Hongyu Xu, Ke Ng, Kam-Woh Pérez, Juan C. Pérez-Rúa, Juan-Manuel Xiang, Tao Liu, Wei Liu, Shikun Schmidhuber, Jürgen |
| author_facet | Liu, Haozhe Liu, Ding Zhuge, Mingchen Zhou, Zijian Xie, Tian He, Sen Yang, Yukang Liu, Shuming Cong, Yuren Guo, Jiadong Xu, Hongyu Xu, Ke Ng, Kam-Woh Pérez, Juan C. Pérez-Rúa, Juan-Manuel Xiang, Tao Liu, Wei Liu, Shikun Schmidhuber, Jürgen |
| contents | We introduce MoS (Mixture of States), a novel fusion paradigm for multimodal diffusion models that merges modalities using flexible, state-based interactions. The core of MoS is a learnable, token-wise router that creates denoising timestep- and input-dependent interactions between modalities' hidden states, precisely aligning token-level features with the diffusion trajectory. This router sparsely selects the top-$k$ hidden states and is trained with an $ε$-greedy strategy, efficiently selecting contextual features with minimal learnable parameters and negligible computational overhead. We validate our design with text-to-image generation (MoS-Image) and editing (MoS-Editing), which achieve state-of-the-art results. With only 3B to 5B parameters, our models match or surpass counterparts up to $4\times$ larger. These findings establish MoS as a flexible and compute-efficient paradigm for scaling multimodal diffusion models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_12207 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Mixture of States: Routing Token-Level Dynamics for Multimodal Generation Liu, Haozhe Liu, Ding Zhuge, Mingchen Zhou, Zijian Xie, Tian He, Sen Yang, Yukang Liu, Shuming Cong, Yuren Guo, Jiadong Xu, Hongyu Xu, Ke Ng, Kam-Woh Pérez, Juan C. Pérez-Rúa, Juan-Manuel Xiang, Tao Liu, Wei Liu, Shikun Schmidhuber, Jürgen Computer Vision and Pattern Recognition We introduce MoS (Mixture of States), a novel fusion paradigm for multimodal diffusion models that merges modalities using flexible, state-based interactions. The core of MoS is a learnable, token-wise router that creates denoising timestep- and input-dependent interactions between modalities' hidden states, precisely aligning token-level features with the diffusion trajectory. This router sparsely selects the top-$k$ hidden states and is trained with an $ε$-greedy strategy, efficiently selecting contextual features with minimal learnable parameters and negligible computational overhead. We validate our design with text-to-image generation (MoS-Image) and editing (MoS-Editing), which achieve state-of-the-art results. With only 3B to 5B parameters, our models match or surpass counterparts up to $4\times$ larger. These findings establish MoS as a flexible and compute-efficient paradigm for scaling multimodal diffusion models. |
| title | Mixture of States: Routing Token-Level Dynamics for Multimodal Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.12207 |