Ming-Omni: A Unified Multimodal Model for Perception and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | AI, Inclusion, Gong, Biao, Zou, Cheng, Zheng, Chuanyang, Zhou, Chunluan, Yan, Canxiang, Jin, Chunxiang, Shen, Chunjie, Zheng, Dandan, Wang, Fudong, Xu, Furong, Yao, GuangMing, Zhou, Jun, Chen, Jingdong, Sun, Jianxin, Liu, Jiajia, Zhu, Jianjiang, Peng, Jun, Ji, Kaixiang, Song, Kaiyou, Ren, Kaimeng, Wang, Libin, Ru, Lixiang, Xie, Lele, Tan, Longhua, Xue, Lyuxin, Wang, Lan, Bai, Mochen, Gao, Ning, Chen, Pei, Guo, Qingpei, Zhang, Qinglong, Xu, Qiang, Liu, Rui, Xiong, Ruijie, Gao, Sirui, Liu, Tinghao, Li, Taisong, Chai, Weilong, Xiao, Xinyu, Wang, Xiaomei, Chen, Xiaoxue, Lu, Xiao, Li, Xiaoyu, Dong, Xingning, Yu, Xuzheng, Yuan, Yi, Gao, Yuting, Sun, Yunxiao, Chen, Yipeng, Wu, Yifei, Lyu, Yongjie, Ma, Ziping, Feng, Zipeng, Fang, Zhijiang, Qiu, Zhihao, Huang, Ziyuan, He, Zhengyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
by: Dong, Xingning, et al.
Published: (2024)
by: Dong, Xingning, et al.
Published: (2024)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
by: Yu, Xuzheng, et al.
Published: (2024)
by: Yu, Xuzheng, et al.
Published: (2024)
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
by: Huang, Ziyuan, et al.
Published: (2025)
by: Huang, Ziyuan, et al.
Published: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
by: Yan, Canxiang, et al.
Published: (2025)
by: Yan, Canxiang, et al.
Published: (2025)
M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
by: Guo, Qingpei, et al.
Published: (2025)
by: Guo, Qingpei, et al.
Published: (2025)
Degrees of freedom of quadratic scalar-nonmetricity theory
by: Chen, Jia-Jun, et al.
Published: (2025)
by: Chen, Jia-Jun, et al.
Published: (2025)
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
by: Gao, Yuting, et al.
Published: (2025)
by: Gao, Yuting, et al.
Published: (2025)
SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories
by: Zhu, Muzhi, et al.
Published: (2025)
by: Zhu, Muzhi, et al.
Published: (2025)
Probing a regular black hole within asymptotically safe gravity via strong gravitational lensings and optical appearances
by: Gao, Xiao-Jun
Published: (2024)
by: Gao, Xiao-Jun
Published: (2024)
Gravitational lensing and shadow by a Schwarzschild-like black hole in metric-affine bumblebee gravity
by: Gao, Xiao-Jun
Published: (2024)
by: Gao, Xiao-Jun
Published: (2024)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
iFormer: Integrating ConvNet and Transformer for Mobile Application
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
Spatially Confined Alloying of Pt Accelerates Mass Transport for Fuel Cell Oxygen Reduction
by: Yuxin Gao, et al.
Published: (2024)
by: Yuxin Gao, et al.
Published: (2024)
An Isolable One‐Coordinate Lead(I) Radical with Strong g‐Factor Anisotropy
by: Haonan Chen, et al.
Published: (2024)
by: Haonan Chen, et al.
Published: (2024)
An Isolable One‐Coordinate Lead(I) Radical with Strong g‐Factor Anisotropy
by: Haonan Chen, et al.
Published: (2024)
by: Haonan Chen, et al.
Published: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
by: Gao, Yuting, et al.
Published: (2025)
by: Gao, Yuting, et al.
Published: (2025)
Boundary Perturbation Effects in Quantum Systems with Conserved Energy and Continuous Symmetry
by: Gao, Qucheng, et al.
Published: (2025)
by: Gao, Qucheng, et al.
Published: (2025)
LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models
by: Zhao, Hengyuan, et al.
Published: (2025)
by: Zhao, Hengyuan, et al.
Published: (2025)
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
by: Gao, Kaifeng, et al.
Published: (2024)
by: Gao, Kaifeng, et al.
Published: (2024)
$\text{Di}^2\text{Pose}$: Discrete Diffusion Model for Occluded 3D Human Pose Estimation
by: Wang, Weiquan, et al.
Published: (2024)
by: Wang, Weiquan, et al.
Published: (2024)
Investigating shadow of a rotating charged black hole with a cosmological constant immersed in the perfect fluid dark matter
by: Ban, Zheng-Long, et al.
Published: (2025)
by: Ban, Zheng-Long, et al.
Published: (2025)
Towards Securing UAV‐Assisted Edge Computing: A Trust‐Based Intrusion Detection Framework With Multi‐Source Feedback
by: Jun Tao, et al.
Published: (2025)
by: Jun Tao, et al.
Published: (2025)
Event-Customized Image Generation
by: Wang, Zhen, et al.
Published: (2024)
by: Wang, Zhen, et al.
Published: (2024)
Semantic Enhanced Few-shot Object Detection
by: Wang, Zheng, et al.
Published: (2024)
by: Wang, Zheng, et al.
Published: (2024)
Heteroatom‐Doped Graphene Nanoribbons: Precision Synthesis and Emerging Properties†
by: Pei‐Han Gao, et al.
Published: (2024)
by: Pei‐Han Gao, et al.
Published: (2024)
Fuel-Optimal Trajectory Planning for Lunar Vertical Landing
by: Wang, Kun, et al.
Published: (2024)
by: Wang, Kun, et al.
Published: (2024)
Fuel-optimal powered descent guidance for lunar pinpoint landing using neural networks
by: Wang, Kun, et al.
Published: (2024)
by: Wang, Kun, et al.
Published: (2024)
Data-Dependent Stability Analysis of Adversarial Training
by: Wang, Yihan, et al.
Published: (2024)
by: Wang, Yihan, et al.
Published: (2024)
Producing $Λ(1405)$ and $Λ(1520)$ in $π^-p$ reaction to explore their inner structures
by: Gao, Yuan, et al.
Published: (2026)
by: Gao, Yuan, et al.
Published: (2026)
Game-Theoretic Unlearnable Example Generator
by: Liu, Shuang, et al.
Published: (2024)
by: Liu, Shuang, et al.
Published: (2024)
Production potential of hidden-strange molecular pentaquarks through the $π^-p\rightarrow K^{*}Σ$ process
by: Wang, Xiao-Yun, et al.
Published: (2024)
by: Wang, Xiao-Yun, et al.
Published: (2024)
Understanding of the BESIII measurement of (anti)hyperon-nucleon scattering
by: Wang, Xiao-Yun, et al.
Published: (2024)
by: Wang, Xiao-Yun, et al.
Published: (2024)
Gravitational Lensing of Spherically Symmetric Black Holes in Dark Matter Halos
by: Liu, Yi-Gao, et al.
Published: (2023)
by: Liu, Yi-Gao, et al.
Published: (2023)
Towards Differentiable Multilevel Optimization: A Gradient-Based Approach
by: Gu, Yuntian, et al.
Published: (2024)
by: Gu, Yuntian, et al.
Published: (2024)
Status of Nano-ARPES endstation at BL07U of Shanghai Synchrotron Radiation Facility
by: Gao, Han, et al.
Published: (2024)
by: Gao, Han, et al.
Published: (2024)
The B‐Site Synergistic Metal Ion Co‐Doping Strategy for Enhancing the Scintillation Performance of 2D Perovskite Single Crystals
by: Zehui Xiang, et al.
Published: (2025)
by: Zehui Xiang, et al.
Published: (2025)
Similar Items
-
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
by: Dong, Xingning, et al.
Published: (2024) -
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
by: AI, Inclusion, et al.
Published: (2025) -
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
by: Yu, Xuzheng, et al.
Published: (2024) -
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
by: AI, Inclusion, et al.
Published: (2025) -
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)