Emu3.5: Native Multimodal Models are World Learners

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Yufeng, Chen, Honghao, Deng, Haoge, Huang, Xu, Li, Xinghang, Liu, Jirong, Liu, Yang, Luo, Zhuoyan, Wang, Jinsheng, Wang, Wenxuan, Wang, Yueze, Wang, Chengyuan, Zhang, Fan, Zhao, Yingli, Pan, Ting, Li, Xianduo, Hao, Zecheng, Ma, Wenxuan, Chen, Zhuo, Ao, Yulong, Huang, Tiejun, Wang, Zhongyuan, Wang, Xinlong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912678350422016
author Cui, Yufeng
Chen, Honghao
Deng, Haoge
Huang, Xu
Li, Xinghang
Liu, Jirong
Liu, Yang
Luo, Zhuoyan
Wang, Jinsheng
Wang, Wenxuan
Wang, Yueze
Wang, Chengyuan
Zhang, Fan
Zhao, Yingli
Pan, Ting
Li, Xianduo
Hao, Zecheng
Ma, Wenxuan
Chen, Zhuo
Ao, Yulong
Huang, Tiejun
Wang, Zhongyuan
Wang, Xinlong
author_facet Cui, Yufeng
Chen, Honghao
Deng, Haoge
Huang, Xu
Li, Xinghang
Liu, Jirong
Liu, Yang
Luo, Zhuoyan
Wang, Jinsheng
Wang, Wenxuan
Wang, Yueze
Wang, Chengyuan
Zhang, Fan
Zhao, Yingli
Pan, Ting
Li, Xianduo
Hao, Zecheng
Ma, Wenxuan
Chen, Zhuo
Ao, Yulong
Huang, Tiejun
Wang, Zhongyuan
Wang, Xinlong
contents We introduce Emu3.5, a large-scale multimodal world model that natively predicts the next state across vision and language. Emu3.5 is pre-trained end-to-end with a unified next-token prediction objective on a corpus of vision-language interleaved data containing over 10 trillion tokens, primarily derived from sequential frames and transcripts of internet videos. The model naturally accepts interleaved vision-language inputs and generates interleaved vision-language outputs. Emu3.5 is further post-trained with large-scale reinforcement learning to enhance multimodal reasoning and generation. To improve inference efficiency, we propose Discrete Diffusion Adaptation (DiDA), which converts token-by-token decoding into bidirectional parallel prediction, accelerating per-image inference by about 20x without sacrificing performance. Emu3.5 exhibits strong native multimodal capabilities, including long-horizon vision-language generation, any-to-image (X2I) generation, and complex text-rich image generation. It also exhibits generalizable world-modeling abilities, enabling spatiotemporally consistent world exploration and open-world embodied manipulation across diverse scenarios and tasks. For comparison, Emu3.5 achieves performance comparable to Gemini 2.5 Flash Image (Nano Banana) on image generation and editing tasks and demonstrates superior results on a suite of interleaved generation tasks. We open-source Emu3.5 at https://github.com/baaivision/Emu3.5 to support community research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26583
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emu3.5: Native Multimodal Models are World Learners
Cui, Yufeng
Chen, Honghao
Deng, Haoge
Huang, Xu
Li, Xinghang
Liu, Jirong
Liu, Yang
Luo, Zhuoyan
Wang, Jinsheng
Wang, Wenxuan
Wang, Yueze
Wang, Chengyuan
Zhang, Fan
Zhao, Yingli
Pan, Ting
Li, Xianduo
Hao, Zecheng
Ma, Wenxuan
Chen, Zhuo
Ao, Yulong
Huang, Tiejun
Wang, Zhongyuan
Wang, Xinlong
Computer Vision and Pattern Recognition
We introduce Emu3.5, a large-scale multimodal world model that natively predicts the next state across vision and language. Emu3.5 is pre-trained end-to-end with a unified next-token prediction objective on a corpus of vision-language interleaved data containing over 10 trillion tokens, primarily derived from sequential frames and transcripts of internet videos. The model naturally accepts interleaved vision-language inputs and generates interleaved vision-language outputs. Emu3.5 is further post-trained with large-scale reinforcement learning to enhance multimodal reasoning and generation. To improve inference efficiency, we propose Discrete Diffusion Adaptation (DiDA), which converts token-by-token decoding into bidirectional parallel prediction, accelerating per-image inference by about 20x without sacrificing performance. Emu3.5 exhibits strong native multimodal capabilities, including long-horizon vision-language generation, any-to-image (X2I) generation, and complex text-rich image generation. It also exhibits generalizable world-modeling abilities, enabling spatiotemporally consistent world exploration and open-world embodied manipulation across diverse scenarios and tasks. For comparison, Emu3.5 achieves performance comparable to Gemini 2.5 Flash Image (Nano Banana) on image generation and editing tasks and demonstrates superior results on a suite of interleaved generation tasks. We open-source Emu3.5 at https://github.com/baaivision/Emu3.5 to support community research.
title Emu3.5: Native Multimodal Models are World Learners
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.26583