Context Unrolling in Omni Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Ceyuan, Lin, Zhijie, Zhao, Yang, Xiao, Fei, He, Hao, Zhao, Qi, Deng, Chaorui, Li, Kunchang, Ding, Zihan, Guo, Yuwei, Wang, Fuyun, Zhu, Fangqi, Nie, Xiaonan, Zhu, Shenhan, Lin, Shanchuan, Li, Hongsheng, Huang, Weilin, Shi, Guang, Fan, Haoqi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918464643399680
author Yang, Ceyuan
Lin, Zhijie
Zhao, Yang
Xiao, Fei
He, Hao
Zhao, Qi
Deng, Chaorui
Li, Kunchang
Ding, Zihan
Guo, Yuwei
Wang, Fuyun
Zhu, Fangqi
Nie, Xiaonan
Zhu, Shenhan
Lin, Shanchuan
Li, Hongsheng
Huang, Weilin
Shi, Guang
Fan, Haoqi
author_facet Yang, Ceyuan
Lin, Zhijie
Zhao, Yang
Xiao, Fei
He, Hao
Zhao, Qi
Deng, Chaorui
Li, Kunchang
Ding, Zihan
Guo, Yuwei
Wang, Fuyun
Zhu, Fangqi
Nie, Xiaonan
Zhu, Shenhan
Lin, Shanchuan
Li, Hongsheng
Huang, Weilin
Shi, Guang
Fan, Haoqi
contents We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21921
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Context Unrolling in Omni Models
Yang, Ceyuan
Lin, Zhijie
Zhao, Yang
Xiao, Fei
He, Hao
Zhao, Qi
Deng, Chaorui
Li, Kunchang
Ding, Zihan
Guo, Yuwei
Wang, Fuyun
Zhu, Fangqi
Nie, Xiaonan
Zhu, Shenhan
Lin, Shanchuan
Li, Hongsheng
Huang, Weilin
Shi, Guang
Fan, Haoqi
Computer Vision and Pattern Recognition
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.
title Context Unrolling in Omni Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.21921