_version_ 1866915954266472448
author Mao, Chaojie
Xie, Chen-Wei
Zhong, Chongyang
Deng, Haoyou
Zhao, Jiaxing
Xiao, Jie
Xing, Jinbo
Zhang, Jingfeng
Zhou, Jingren
Zhang, Jingyi
Dan, Jun
Zhu, Kai
Zhao, Kang
Yan, Keyu
Chen, Minghui
Li, Pandeng
Chen, Shuangle
Shen, Tong
Liu, Yu
Jiang, Yue
Pan, Yulin
Tuo, Yuxiang
Jiang, Zeyinzi
Han, Zhen
Wang, Ang
Zhang, Bang
Ai, Baole
Wen, Bin
Feng, Boang
Yu, Feiwu
Wang, Gang
Zhao, Haiming
Kang, He
Xiang, Jianjing
Zeng, Jianyuan
Wang, Jinkai
Zhou, Junjie
Sun, Ke
Wu, Linqian
Gong, Pei
Wu, Pingyu
Wu, Ruiwen
Su, Tongtong
Zhou, Wenmeng
Shen, Wenting
Yu, Wenyuan
Xu, Xianjun
Huang, Xiaoming
Shen, Xiejie
Xu, Xin
Kou, Yan
Lv, Yangyu
Zhai, Yifan
Huang, Yitong
Zheng, Yun
Hong, Yuntao
Zhang, Zhe
Zhang, Zhicheng
author_facet Mao, Chaojie
Xie, Chen-Wei
Zhong, Chongyang
Deng, Haoyou
Zhao, Jiaxing
Xiao, Jie
Xing, Jinbo
Zhang, Jingfeng
Zhou, Jingren
Zhang, Jingyi
Dan, Jun
Zhu, Kai
Zhao, Kang
Yan, Keyu
Chen, Minghui
Li, Pandeng
Chen, Shuangle
Shen, Tong
Liu, Yu
Jiang, Yue
Pan, Yulin
Tuo, Yuxiang
Jiang, Zeyinzi
Han, Zhen
Wang, Ang
Zhang, Bang
Ai, Baole
Wen, Bin
Feng, Boang
Yu, Feiwu
Wang, Gang
Zhao, Haiming
Kang, He
Xiang, Jianjing
Zeng, Jianyuan
Wang, Jinkai
Zhou, Junjie
Sun, Ke
Wu, Linqian
Gong, Pei
Wu, Pingyu
Wu, Ruiwen
Su, Tongtong
Zhou, Wenmeng
Shen, Wenting
Yu, Wenyuan
Xu, Xianjun
Huang, Xiaoming
Shen, Xiejie
Xu, Xin
Kou, Yan
Lv, Yangyu
Zhai, Yifan
Huang, Yitong
Zheng, Yun
Hong, Yuntao
Zhang, Zhe
Zhang, Zhicheng
contents We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering, and strict identity preservation. To address these challenges, Wan-Image features a natively unified multi-modal architecture by synergizing the cognitive capabilities of large language models with the high-fidelity pixel synthesis of diffusion transformers, which seamlessly translates highly nuanced user intents into precise visual outputs. It is fundamentally powered by large-scale multi-modal data scaling, a systematic fine-grained annotation engine, and curated reinforcement learning data to surpass basic instruction following and unlock expert-level professional capabilities. These include ultra-long complex text rendering, hyper-diverse portrait generation, palette-guided generation, multi-subject identity preservation, coherent sequential visual generation, precise multi-modal interactive editing, native alpha-channel generation, and high-efficiency 4K synthesis. Across diverse human evaluations, Wan-Image exceeds Seedream 5.0 Lite and GPT Image 1.5 in overall performance, reaching parity with Nano Banana Pro in challenging tasks. Ultimately, Wan-Image revolutionizes visual content creation across e-commerce, entertainment, education, and personal productivity, redefining the boundaries of professional visual synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19858
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
Mao, Chaojie
Xie, Chen-Wei
Zhong, Chongyang
Deng, Haoyou
Zhao, Jiaxing
Xiao, Jie
Xing, Jinbo
Zhang, Jingfeng
Zhou, Jingren
Zhang, Jingyi
Dan, Jun
Zhu, Kai
Zhao, Kang
Yan, Keyu
Chen, Minghui
Li, Pandeng
Chen, Shuangle
Shen, Tong
Liu, Yu
Jiang, Yue
Pan, Yulin
Tuo, Yuxiang
Jiang, Zeyinzi
Han, Zhen
Wang, Ang
Zhang, Bang
Ai, Baole
Wen, Bin
Feng, Boang
Yu, Feiwu
Wang, Gang
Zhao, Haiming
Kang, He
Xiang, Jianjing
Zeng, Jianyuan
Wang, Jinkai
Zhou, Junjie
Sun, Ke
Wu, Linqian
Gong, Pei
Wu, Pingyu
Wu, Ruiwen
Su, Tongtong
Zhou, Wenmeng
Shen, Wenting
Yu, Wenyuan
Xu, Xianjun
Huang, Xiaoming
Shen, Xiejie
Xu, Xin
Kou, Yan
Lv, Yangyu
Zhai, Yifan
Huang, Yitong
Zheng, Yun
Hong, Yuntao
Zhang, Zhe
Zhang, Zhicheng
Computer Vision and Pattern Recognition
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering, and strict identity preservation. To address these challenges, Wan-Image features a natively unified multi-modal architecture by synergizing the cognitive capabilities of large language models with the high-fidelity pixel synthesis of diffusion transformers, which seamlessly translates highly nuanced user intents into precise visual outputs. It is fundamentally powered by large-scale multi-modal data scaling, a systematic fine-grained annotation engine, and curated reinforcement learning data to surpass basic instruction following and unlock expert-level professional capabilities. These include ultra-long complex text rendering, hyper-diverse portrait generation, palette-guided generation, multi-subject identity preservation, coherent sequential visual generation, precise multi-modal interactive editing, native alpha-channel generation, and high-efficiency 4K synthesis. Across diverse human evaluations, Wan-Image exceeds Seedream 5.0 Lite and GPT Image 1.5 in overall performance, reaching parity with Nano Banana Pro in challenging tasks. Ultimately, Wan-Image revolutionizes visual content creation across e-commerce, entertainment, education, and personal productivity, redefining the boundaries of professional visual synthesis.
title Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.19858