Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917361982898176 |
|---|---|
| author | AI, Inclusion : Ma, Bowen Zou, Cheng Du, ChengKun Yan, Canxiang Jin, Chunxiang Shen, Chunjie Lian, Chenyu Fan, Chengxiang Zheng, Dandan Wang, Fudong Xu, Furong Yao, Guangming Liu, Haohao Peng, Han Zhou, Jun Xia, Junluan Chen, Jingdong Li, Jianing Sun, Jianxin Zhu, Jianjiang Jiang, Jianping Ou, Jinpeng Peng, Jun Peng, Jin Ji, Kaixiang Tang, Li Wang, Libin Ru, Lixiang Tan, Longhua Ma, Lu Wang, Lan Bai, Mochen Cai, Minghong Yang, Mingxue Gao, Ning Guo, Qingpei Zhang, Qinglong Xu, Qiang Zhao, Qin Liu, Rui Xiong, Ruijie Zheng, Ruobing Gao, Sirui Lin, Shaoxiong Zhang, Tao Li, Tianqi Liu, Tinghao Wang, Tongli Huang, Taoye Chai, Weilong Wang, Xiaomei Wang, Xiaolong Liu, Xiaojian Lu, Xiao Li, Xiaoyu Dong, Xingning Yu, Xuzheng Wang, Xuezhi Yuan, Yi Gao, Yuting Xiao, Yuting Sun, Yunxiao Chen, Yipeng Mao, Yifan Wu, Yifei Lyu, Yongjie Zhang, Yingying Li, YuQian Ma, Ziping Fang, Zhiqiang Qiu, Zhihao Huang, Ziyuan Yang, Zizheng He, Zhengyu |
| author_facet | AI, Inclusion : Ma, Bowen Zou, Cheng Du, ChengKun Yan, Canxiang Jin, Chunxiang Shen, Chunjie Lian, Chenyu Fan, Chengxiang Zheng, Dandan Wang, Fudong Xu, Furong Yao, Guangming Liu, Haohao Peng, Han Zhou, Jun Xia, Junluan Chen, Jingdong Li, Jianing Sun, Jianxin Zhu, Jianjiang Jiang, Jianping Ou, Jinpeng Peng, Jun Peng, Jin Ji, Kaixiang Tang, Li Wang, Libin Ru, Lixiang Tan, Longhua Ma, Lu Wang, Lan Bai, Mochen Cai, Minghong Yang, Mingxue Gao, Ning Guo, Qingpei Zhang, Qinglong Xu, Qiang Zhao, Qin Liu, Rui Xiong, Ruijie Zheng, Ruobing Gao, Sirui Lin, Shaoxiong Zhang, Tao Li, Tianqi Liu, Tinghao Wang, Tongli Huang, Taoye Chai, Weilong Wang, Xiaomei Wang, Xiaolong Liu, Xiaojian Lu, Xiao Li, Xiaoyu Dong, Xingning Yu, Xuzheng Wang, Xuezhi Yuan, Yi Gao, Yuting Xiao, Yuting Sun, Yunxiao Chen, Yipeng Mao, Yifan Wu, Yifei Lyu, Yongjie Zhang, Yingying Li, YuQian Ma, Ziping Fang, Zhiqiang Qiu, Zhihao Huang, Ziyuan Yang, Zizheng He, Zhengyu |
| contents | We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_24821 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation AI, Inclusion : Ma, Bowen Zou, Cheng Du, ChengKun Yan, Canxiang Jin, Chunxiang Shen, Chunjie Lian, Chenyu Fan, Chengxiang Zheng, Dandan Wang, Fudong Xu, Furong Yao, Guangming Liu, Haohao Peng, Han Zhou, Jun Xia, Junluan Chen, Jingdong Li, Jianing Sun, Jianxin Zhu, Jianjiang Jiang, Jianping Ou, Jinpeng Peng, Jun Peng, Jin Ji, Kaixiang Tang, Li Wang, Libin Ru, Lixiang Tan, Longhua Ma, Lu Wang, Lan Bai, Mochen Cai, Minghong Yang, Mingxue Gao, Ning Guo, Qingpei Zhang, Qinglong Xu, Qiang Zhao, Qin Liu, Rui Xiong, Ruijie Zheng, Ruobing Gao, Sirui Lin, Shaoxiong Zhang, Tao Li, Tianqi Liu, Tinghao Wang, Tongli Huang, Taoye Chai, Weilong Wang, Xiaomei Wang, Xiaolong Liu, Xiaojian Lu, Xiao Li, Xiaoyu Dong, Xingning Yu, Xuzheng Wang, Xuezhi Yuan, Yi Gao, Yuting Xiao, Yuting Sun, Yunxiao Chen, Yipeng Mao, Yifan Wu, Yifei Lyu, Yongjie Zhang, Yingying Li, YuQian Ma, Ziping Fang, Zhiqiang Qiu, Zhihao Huang, Ziyuan Yang, Zizheng He, Zhengyu Computer Vision and Pattern Recognition Artificial Intelligence We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence. |
| title | Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2510.24821 |