Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913890698264576 |
|---|---|
| author | AI, Inclusion Gong, Biao Zou, Cheng Zheng, Dandan Yu, Hu Chen, Jingdong Sun, Jianxin Zhao, Junbo Zhou, Jun Ji, Kaixiang Ru, Lixiang Wang, Libin Guo, Qingpei Liu, Rui Chai, Weilong Xiao, Xinyu Huang, Ziyuan |
| author_facet | AI, Inclusion Gong, Biao Zou, Cheng Zheng, Dandan Yu, Hu Chen, Jingdong Sun, Jianxin Zhao, Junbo Zhou, Jun Ji, Kaixiang Ru, Lixiang Wang, Libin Guo, Qingpei Liu, Rui Chai, Weilong Xiao, Xinyu Huang, Ziyuan |
| contents | We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides an open-source implementation of the integrated MetaQueries and M2-omni framework, while introducing the novel multi-scale learnable tokens and multi-scale representation alignment strategy. By leveraging a fixed MLLM and a learnable diffusion model, Ming-Lite-Uni enables native multimodal AR models to perform both text-to-image generation and instruction based image editing tasks, expanding their capabilities beyond pure visual understanding. Our experimental results demonstrate the strong performance of Ming-Lite-Uni and illustrate the impressive fluid nature of its interactive process. All code and model weights are open-sourced to foster further exploration within the community. Notably, this work aligns with concurrent multimodal AI milestones - such as ChatGPT-4o with native image generation updated in March 25, 2025 - underscoring the broader significance of unified models like Ming-Lite-Uni on the path toward AGI. Ming-Lite-Uni is in alpha stage and will soon be further refined. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_02471 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction AI, Inclusion Gong, Biao Zou, Cheng Zheng, Dandan Yu, Hu Chen, Jingdong Sun, Jianxin Zhao, Junbo Zhou, Jun Ji, Kaixiang Ru, Lixiang Wang, Libin Guo, Qingpei Liu, Rui Chai, Weilong Xiao, Xinyu Huang, Ziyuan Computer Vision and Pattern Recognition We introduce Ming-Lite-Uni, an open-source multimodal framework featuring a newly designed unified visual generator and a native multimodal autoregressive model tailored for unifying vision and language. Specifically, this project provides an open-source implementation of the integrated MetaQueries and M2-omni framework, while introducing the novel multi-scale learnable tokens and multi-scale representation alignment strategy. By leveraging a fixed MLLM and a learnable diffusion model, Ming-Lite-Uni enables native multimodal AR models to perform both text-to-image generation and instruction based image editing tasks, expanding their capabilities beyond pure visual understanding. Our experimental results demonstrate the strong performance of Ming-Lite-Uni and illustrate the impressive fluid nature of its interactive process. All code and model weights are open-sourced to foster further exploration within the community. Notably, this work aligns with concurrent multimodal AI milestones - such as ChatGPT-4o with native image generation updated in March 25, 2025 - underscoring the broader significance of unified models like Ming-Lite-Uni on the path toward AGI. Ming-Lite-Uni is in alpha stage and will soon be further refined. |
| title | Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2505.02471 |