Saved in:
Bibliographic Details
Main Authors: Yang, Ceyuan, Lin, Zhijie, Zhao, Yang, Xiao, Fei, He, Hao, Zhao, Qi, Deng, Chaorui, Li, Kunchang, Ding, Zihan, Guo, Yuwei, Wang, Fuyun, Zhu, Fangqi, Nie, Xiaonan, Zhu, Shenhan, Lin, Shanchuan, Li, Hongsheng, Huang, Weilin, Shi, Guang, Fan, Haoqi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.21921
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918464643399680
author Yang, Ceyuan
Lin, Zhijie
Zhao, Yang
Xiao, Fei
He, Hao
Zhao, Qi
Deng, Chaorui
Li, Kunchang
Ding, Zihan
Guo, Yuwei
Wang, Fuyun
Zhu, Fangqi
Nie, Xiaonan
Zhu, Shenhan
Lin, Shanchuan
Li, Hongsheng
Huang, Weilin
Shi, Guang
Fan, Haoqi
author_facet Yang, Ceyuan
Lin, Zhijie
Zhao, Yang
Xiao, Fei
He, Hao
Zhao, Qi
Deng, Chaorui
Li, Kunchang
Ding, Zihan
Guo, Yuwei
Wang, Fuyun
Zhu, Fangqi
Nie, Xiaonan
Zhu, Shenhan
Lin, Shanchuan
Li, Hongsheng
Huang, Weilin
Shi, Guang
Fan, Haoqi
contents We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21921
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Context Unrolling in Omni Models
Yang, Ceyuan
Lin, Zhijie
Zhao, Yang
Xiao, Fei
He, Hao
Zhao, Qi
Deng, Chaorui
Li, Kunchang
Ding, Zihan
Guo, Yuwei
Wang, Fuyun
Zhu, Fangqi
Nie, Xiaonan
Zhu, Shenhan
Lin, Shanchuan
Li, Hongsheng
Huang, Weilin
Shi, Guang
Fan, Haoqi
Computer Vision and Pattern Recognition
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.
title Context Unrolling in Omni Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.21921