Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jaeik, Kim, Woojin, Hong, Jihwan, Lee, Yejoon, Hyeon, Sieun, Lim, Mintaek, Han, Yunseok, Kim, Dogeun, Lee, Hoeun, Kim, Hyunggeun, Do, Jaeyoung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917375766429696
author Kim, Jaeik
Kim, Woojin
Hong, Jihwan
Lee, Yejoon
Hyeon, Sieun
Lim, Mintaek
Han, Yunseok
Kim, Dogeun
Lee, Hoeun
Kim, Hyunggeun
Do, Jaeyoung
author_facet Kim, Jaeik
Kim, Woojin
Hong, Jihwan
Lee, Yejoon
Hyeon, Sieun
Lim, Mintaek
Han, Yunseok
Kim, Dogeun
Lee, Hoeun
Kim, Hyunggeun
Do, Jaeyoung
contents We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive unified models that serialize heterogeneous modalities, or compositional unified models that require orchestration with external modality-specific decoders, Dynin-Omni natively formulates omnimodal modeling as masked diffusion over a shared discrete token space, enabling iterative refinement under bidirectional context. Dynin-Omni adopts a multi-stage training strategy with model-merging-based modality expansion and omnimodal alignment. We evaluate Dynin-Omni across 19 multimodal benchmarks spanning language reasoning, image generation and editing, video understanding, and speech recognition and synthesis. Dynin-Omni achieves 87.6 on GSM8K, 1733.6 on MME-P, 61.4 on VideoMME, 0.87 on GenEval, and 2.1 WER on LibriSpeech test-clean, consistently outperforming existing open-source unified models while remaining competitive with strong modality-specific expert systems. These results demonstrate the potential of masked diffusion as a unified paradigm for any-to-any modeling, providing a flexible foundation for real-time omnimodal systems, unified cross-modal retrieval and generation, and embodied multimodal agents.
format Preprint
id arxiv_https___arxiv_org_abs_2604_00007
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dynin-Omni: Omnimodal Unified Large Diffusion Language Model
Kim, Jaeik
Kim, Woojin
Hong, Jihwan
Lee, Yejoon
Hyeon, Sieun
Lim, Mintaek
Han, Yunseok
Kim, Dogeun
Lee, Hoeun
Kim, Hyunggeun
Do, Jaeyoung
Computation and Language
Artificial Intelligence
We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlike autoregressive unified models that serialize heterogeneous modalities, or compositional unified models that require orchestration with external modality-specific decoders, Dynin-Omni natively formulates omnimodal modeling as masked diffusion over a shared discrete token space, enabling iterative refinement under bidirectional context. Dynin-Omni adopts a multi-stage training strategy with model-merging-based modality expansion and omnimodal alignment. We evaluate Dynin-Omni across 19 multimodal benchmarks spanning language reasoning, image generation and editing, video understanding, and speech recognition and synthesis. Dynin-Omni achieves 87.6 on GSM8K, 1733.6 on MME-P, 61.4 on VideoMME, 0.87 on GenEval, and 2.1 WER on LibriSpeech test-clean, consistently outperforming existing open-source unified models while remaining competitive with strong modality-specific expert systems. These results demonstrate the potential of masked diffusion as a unified paradigm for any-to-any modeling, providing a flexible foundation for real-time omnimodal systems, unified cross-modal retrieval and generation, and embodied multimodal agents.
title Dynin-Omni: Omnimodal Unified Large Diffusion Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.00007