Dynin-Omni: Omnimodal Unified Large Diffusion Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917375766429696 |
|---|---|
| author | Kim, Jaeik Kim, Woojin Hong, Jihwan Lee, Yejoon Hyeon, Sieun Lim, Mintaek Han, Yunseok Kim, Dogeun Lee, Hoeun Kim, Hyunggeun Do, Jaeyoung |
| author_facet | Kim, Jaeik Kim, Woojin Hong, Jihwan Lee, Yejoon Hyeon, Sieun Lim, Mintaek Han, Yunseok Kim, Dogeun Lee, Hoeun Kim, Hyunggeun Do, Jaeyoung |
| contents | We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive unified models that serialize heterogeneous modalities, or compositional unified models that require orchestration with external modality-specific decoders, Dynin-Omni natively formulates omnimodal modeling as masked diffusion over a shared discrete token space, enabling iterative refinement under bidirectional context. Dynin-Omni adopts a multi-stage training strategy with model-merging-based modality expansion and omnimodal alignment. We evaluate Dynin-Omni across 19 multimodal benchmarks spanning language reasoning, image generation and editing, video understanding, and speech recognition and synthesis. Dynin-Omni achieves 87.6 on GSM8K, 1733.6 on MME-P, 61.4 on VideoMME, 0.87 on GenEval, and 2.1 WER on LibriSpeech test-clean, consistently outperforming existing open-source unified models while remaining competitive with strong modality-specific expert systems. These results demonstrate the potential of masked diffusion as a unified paradigm for any-to-any modeling, providing a flexible foundation for real-time omnimodal systems, unified cross-modal retrieval and generation, and embodied multimodal agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_00007 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Dynin-Omni: Omnimodal Unified Large Diffusion Language Model Kim, Jaeik Kim, Woojin Hong, Jihwan Lee, Yejoon Hyeon, Sieun Lim, Mintaek Han, Yunseok Kim, Dogeun Lee, Hoeun Kim, Hyunggeun Do, Jaeyoung Computation and Language Artificial Intelligence We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive unified models that serialize heterogeneous modalities, or compositional unified models that require orchestration with external modality-specific decoders, Dynin-Omni natively formulates omnimodal modeling as masked diffusion over a shared discrete token space, enabling iterative refinement under bidirectional context. Dynin-Omni adopts a multi-stage training strategy with model-merging-based modality expansion and omnimodal alignment. We evaluate Dynin-Omni across 19 multimodal benchmarks spanning language reasoning, image generation and editing, video understanding, and speech recognition and synthesis. Dynin-Omni achieves 87.6 on GSM8K, 1733.6 on MME-P, 61.4 on VideoMME, 0.87 on GenEval, and 2.1 WER on LibriSpeech test-clean, consistently outperforming existing open-source unified models while remaining competitive with strong modality-specific expert systems. These results demonstrate the potential of masked diffusion as a unified paradigm for any-to-any modeling, providing a flexible foundation for real-time omnimodal systems, unified cross-modal retrieval and generation, and embodied multimodal agents. |
| title | Dynin-Omni: Omnimodal Unified Large Diffusion Language Model |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.00007 |