AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Dongjie, Yuan, Ruifeng, Li, Yongqi, You, Runyang, Wang, Wenjie, Nie, Liqiang, Zhang, Lei, Li, Wenjie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912846289305600
author Cheng, Dongjie
Yuan, Ruifeng
Li, Yongqi
You, Runyang
Wang, Wenjie
Nie, Liqiang
Zhang, Lei
Li, Wenjie
author_facet Cheng, Dongjie
Yuan, Ruifeng
Li, Yongqi
You, Runyang
Wang, Wenjie
Nie, Liqiang
Zhang, Lei
Li, Wenjie
contents Real-world perception and interaction are inherently multimodal, encompassing not only language but also vision and speech, which motivates the development of "Omni" MLLMs that support both multimodal inputs and multimodal outputs. While a sequence of omni MLLMs has emerged, most existing systems still rely on additional expert components to achieve multimodal generation, limiting the simplicity of unified training and inference. Autoregressive (AR) modeling, with a single token stream, a single next-token objective, and a single decoder, is an elegant and scalable foundation in the text domain. Motivated by this, we present AR-Omni, a unified any-to-any model in the autoregressive paradigm without any expert decoders. AR-Omni supports autoregressive text and image generation, as well as streaming speech generation, all under a single Transformer decoder. We further address three practical issues in unified AR modeling: modality imbalance via task-aware loss reweighting, visual fidelity via a lightweight token-level perceptual alignment loss for image tokens, and stability-creativity trade-offs via a finite-state decoding mechanism. Empirically, AR-Omni achieves strong quality across three modalities while remaining real-time, achieving a 0.88 real-time factor for speech generation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17761
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
Cheng, Dongjie
Yuan, Ruifeng
Li, Yongqi
You, Runyang
Wang, Wenjie
Nie, Liqiang
Zhang, Lei
Li, Wenjie
Machine Learning
Artificial Intelligence
Computation and Language
Real-world perception and interaction are inherently multimodal, encompassing not only language but also vision and speech, which motivates the development of "Omni" MLLMs that support both multimodal inputs and multimodal outputs. While a sequence of omni MLLMs has emerged, most existing systems still rely on additional expert components to achieve multimodal generation, limiting the simplicity of unified training and inference. Autoregressive (AR) modeling, with a single token stream, a single next-token objective, and a single decoder, is an elegant and scalable foundation in the text domain. Motivated by this, we present AR-Omni, a unified any-to-any model in the autoregressive paradigm without any expert decoders. AR-Omni supports autoregressive text and image generation, as well as streaming speech generation, all under a single Transformer decoder. We further address three practical issues in unified AR modeling: modality imbalance via task-aware loss reweighting, visual fidelity via a lightweight token-level perceptual alignment loss for image tokens, and stability-creativity trade-offs via a finite-state decoding mechanism. Empirically, AR-Omni achieves strong quality across three modalities while remaining real-time, achieving a 0.88 real-time factor for speech generation.
title AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.17761