M2-omni: Advancing Omni-MLLM for Comprehensive Modality Support with Competitive Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Qingpei, Song, Kaiyou, Feng, Zipeng, Ma, Ziping, Zhang, Qinglong, Gao, Sirui, Yu, Xuzheng, Sun, Yunxiao, Chang, Tai-Wei, Chen, Jingdong, Yang, Ming, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
von: Dong, Xingning, et al.
Veröffentlicht: (2024)
von: Dong, Xingning, et al.
Veröffentlicht: (2024)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
von: Ma, Ziping, et al.
Veröffentlicht: (2024)
von: Ma, Ziping, et al.
Veröffentlicht: (2024)
SHE-Net: Syntax-Hierarchy-Enhanced Text-Video Retrieval
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
von: Yu, Xuzheng, et al.
Veröffentlicht: (2024)
New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
von: Yang, Xuzheng, et al.
Veröffentlicht: (2025)
von: Yang, Xuzheng, et al.
Veröffentlicht: (2025)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023)
LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models
von: Zhao, Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Hengyuan, et al.
Veröffentlicht: (2025)
Fuente sonora omni-direccional
von: A. Pérez López
Veröffentlicht: (2006)
von: A. Pérez López
Veröffentlicht: (2006)
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
von: Bie, Fuqing, et al.
Veröffentlicht: (2025)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
von: Zhang, Guohui, et al.
Veröffentlicht: (2026)
von: Zhang, Guohui, et al.
Veröffentlicht: (2026)
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
von: Sun, Qixin, et al.
Veröffentlicht: (2025)
von: Sun, Qixin, et al.
Veröffentlicht: (2025)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
Reliability Assessment of Simply Supported Beam Bridge Structural Performance Based on Load Tests
von: Xuzheng Liu, et al.
Veröffentlicht: (2026)
von: Xuzheng Liu, et al.
Veröffentlicht: (2026)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
von: Jiang, Songtao, et al.
Veröffentlicht: (2024)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
von: Jin, Zhuoran, et al.
Veröffentlicht: (2025)
von: Jin, Zhuoran, et al.
Veröffentlicht: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026)
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
Hummer: Towards Limited Competitive Preference Dataset
von: Jiang, Li, et al.
Veröffentlicht: (2024)
von: Jiang, Li, et al.
Veröffentlicht: (2024)
Learning to Tune Like an Expert: Interpretable and Scene-Aware Navigation via MLLM Reasoning and CVAE-Based Adaptation
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
von: Yin, Bo-Wen, et al.
Veröffentlicht: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024)
von: Liu, Junzhuo, et al.
Veröffentlicht: (2024)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
Is Extending Modality The Right Path Towards Omni-Modality?
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
von: Henry, Felix, et al.
Veröffentlicht: (2026)
von: Henry, Felix, et al.
Veröffentlicht: (2026)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
Omni-Mol: Multitask Molecular Model for Any-to-any Modalities
von: Hu, Chengxin, et al.
Veröffentlicht: (2025)
von: Hu, Chengxin, et al.
Veröffentlicht: (2025)
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models
von: Huang, Haohang, et al.
Veröffentlicht: (2026)
von: Huang, Haohang, et al.
Veröffentlicht: (2026)
Kinematics modeling and simulation of an autonomous omni-directional mobile robot
von: D. Garcia-Sillas
Veröffentlicht: (2015)
von: D. Garcia-Sillas
Veröffentlicht: (2015)
Benchmarking MLLM-based Web Understanding: Reasoning, Robustness and Safety
von: Liu, Junliang, et al.
Veröffentlicht: (2025)
von: Liu, Junliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text Retrieval
von: Dong, Xingning, et al.
Veröffentlicht: (2024) -
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
von: Guo, Qingpei, et al.
Veröffentlicht: (2024) -
Ming-Omni: A Unified Multimodal Model for Perception and Generation
von: AI, Inclusion, et al.
Veröffentlicht: (2025) -
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2026) -
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
von: Ma, Ziping, et al.
Veröffentlicht: (2024)