Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Lijiang, Long, Zuwei, Shen, Yunhang, Gao, Heting, Cao, Haoyu, Sun, Xing, Shan, Caifeng, He, Ran, Fu, Chaoyou
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917319092994048
author Li, Lijiang
Long, Zuwei
Shen, Yunhang
Gao, Heting
Cao, Haoyu
Sun, Xing
Shan, Caifeng
He, Ran
Fu, Chaoyou
author_facet Li, Lijiang
Long, Zuwei
Shen, Yunhang
Gao, Heting
Cao, Haoyu
Sun, Xing
Shan, Caifeng
He, Ran
Fu, Chaoyou
contents While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient alternatives in architectural design. Concurrently, recent studies have successfully applied discrete diffusion models to various domains, such as visual understanding and image generation, revealing their considerable potential as a promising backbone for multimodal systems. Drawing inspiration from these pioneering research, we introduce Omni-Diffusion, the first any-to-any multimodal language model built entirely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. Omni-Diffusion employs a unified mask-based discrete diffusion model to directly capture the joint distribution over discrete multimodal tokens. This approach supports not only bimodal tasks but also more complex scenarios involving multiple modalities. On a diverse set of benchmarks, our method outperforms or performs on par with existing multimodal systems that process two or more modalities, highlighting the significant promise of diffusion models in powering the next generation of multimodal foundation models. Project webpage: https://omni-diffusion.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06577
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
Li, Lijiang
Long, Zuwei
Shen, Yunhang
Gao, Heting
Cao, Haoyu
Sun, Xing
Shan, Caifeng
He, Ran
Fu, Chaoyou
Computer Vision and Pattern Recognition
While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient alternatives in architectural design. Concurrently, recent studies have successfully applied discrete diffusion models to various domains, such as visual understanding and image generation, revealing their considerable potential as a promising backbone for multimodal systems. Drawing inspiration from these pioneering research, we introduce Omni-Diffusion, the first any-to-any multimodal language model built entirely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. Omni-Diffusion employs a unified mask-based discrete diffusion model to directly capture the joint distribution over discrete multimodal tokens. This approach supports not only bimodal tasks but also more complex scenarios involving multiple modalities. On a diverse set of benchmarks, our method outperforms or performs on par with existing multimodal systems that process two or more modalities, highlighting the significant promise of diffusion models in powering the next generation of multimodal foundation models. Project webpage: https://omni-diffusion.github.io.
title Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.06577