Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: AI, Inclusion, Gong, Biao, Zou, Cheng, Zheng, Dandan, Yu, Hu, Chen, Jingdong, Sun, Jianxin, Zhao, Junbo, Zhou, Jun, Ji, Kaixiang, Ru, Lixiang, Wang, Libin, Guo, Qingpei, Liu, Rui, Chai, Weilong, Xiao, Xinyu, Huang, Ziyuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913890698264576
author AI, Inclusion
Gong, Biao
Zou, Cheng
Zheng, Dandan
Yu, Hu
Chen, Jingdong
Sun, Jianxin
Zhao, Junbo
Zhou, Jun
Ji, Kaixiang
Ru, Lixiang
Wang, Libin
Guo, Qingpei
Liu, Rui
Chai, Weilong
Xiao, Xinyu
Huang, Ziyuan
author_facet AI, Inclusion
Gong, Biao
Zou, Cheng
Zheng, Dandan
Yu, Hu
Chen, Jingdong
Sun, Jianxin
Zhao, Junbo
Zhou, Jun
Ji, Kaixiang
Ru, Lixiang
Wang, Libin
Guo, Qingpei
Liu, Rui
Chai, Weilong
Xiao, Xinyu
Huang, Ziyuan
contents We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides an open-source implementation of the integrated MetaQueries and M2-omni framework, while introducing the novel multi-scale learnable tokens and multi-scale representation alignment strategy. By leveraging a fixed MLLM and a learnable diffusion model, Ming-Lite-Uni enables native multimodal AR models to perform both text-to-image generation and instruction based image editing tasks, expanding their capabilities beyond pure visual understanding. Our experimental results demonstrate the strong performance of Ming-Lite-Uni and illustrate the impressive fluid nature of its interactive process. All code and model weights are open-sourced to foster further exploration within the community. Notably, this work aligns with concurrent multimodal AI milestones - such as ChatGPT-4o with native image generation updated in March 25, 2025 - underscoring the broader significance of unified models like Ming-Lite-Uni on the path toward AGI. Ming-Lite-Uni is in alpha stage and will soon be further refined.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
AI, Inclusion
Gong, Biao
Zou, Cheng
Zheng, Dandan
Yu, Hu
Chen, Jingdong
Sun, Jianxin
Zhao, Junbo
Zhou, Jun
Ji, Kaixiang
Ru, Lixiang
Wang, Libin
Guo, Qingpei
Liu, Rui
Chai, Weilong
Xiao, Xinyu
Huang, Ziyuan
Computer Vision and Pattern Recognition
We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides an open-source implementation of the integrated MetaQueries and M2-omni framework, while introducing the novel multi-scale learnable tokens and multi-scale representation alignment strategy. By leveraging a fixed MLLM and a learnable diffusion model, Ming-Lite-Uni enables native multimodal AR models to perform both text-to-image generation and instruction based image editing tasks, expanding their capabilities beyond pure visual understanding. Our experimental results demonstrate the strong performance of Ming-Lite-Uni and illustrate the impressive fluid nature of its interactive process. All code and model weights are open-sourced to foster further exploration within the community. Notably, this work aligns with concurrent multimodal AI milestones - such as ChatGPT-4o with native image generation updated in March 25, 2025 - underscoring the broader significance of unified models like Ming-Lite-Uni on the path toward AGI. Ming-Lite-Uni is in alpha stage and will soon be further refined.
title Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.02471