M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | AI, Inclusion, :, Wang, Fudong, Liu, Jiajia, Chen, Jingdong, Zhou, Jun, Ji, Kaixiang, Ru, Lixiang, Guo, Qingpei, Zheng, Ruobing, Li, Tianqi, Yuan, Yi, Mao, Yifan, Xiao, Yuting, Ma, Ziping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026)
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026)
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
von: Ou, Linyu, et al.
Veröffentlicht: (2025)
von: Ou, Linyu, et al.
Veröffentlicht: (2025)
MM-THEBench: Do Reasoning MLLMs Think Reasonably?
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
von: Huang, Zhidian, et al.
Veröffentlicht: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025)
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
von: Zhao, Qingfei, et al.
Veröffentlicht: (2025)
von: Zhao, Qingfei, et al.
Veröffentlicht: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
von: Fang, I-Sheng, et al.
Veröffentlicht: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
von: Huang, Jincai, et al.
Veröffentlicht: (2026)
von: Huang, Jincai, et al.
Veröffentlicht: (2026)
SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
von: Zhang, Yingying, et al.
Veröffentlicht: (2025)
von: Zhang, Yingying, et al.
Veröffentlicht: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
von: Ru, Jinghan, et al.
Veröffentlicht: (2026)
von: Ru, Jinghan, et al.
Veröffentlicht: (2026)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
von: Ji, Yikun, et al.
Veröffentlicht: (2025)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
von: Wang, Qi, et al.
Veröffentlicht: (2025)
von: Wang, Qi, et al.
Veröffentlicht: (2025)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
von: Chen, Zhenghao, et al.
Veröffentlicht: (2026)
von: Chen, Zhenghao, et al.
Veröffentlicht: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
von: Zhang, Yuting, et al.
Veröffentlicht: (2025)
von: Zhang, Yuting, et al.
Veröffentlicht: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Huanyu, et al.
Veröffentlicht: (2025)
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
von: Yuan, Jiakang, et al.
Veröffentlicht: (2025)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
von: Li, Tianqi, et al.
Veröffentlicht: (2024)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
von: Zhu, Fangrui, et al.
Veröffentlicht: (2025)
von: Zhu, Fangrui, et al.
Veröffentlicht: (2025)
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
von: Wang, Haicheng, et al.
Veröffentlicht: (2026)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
von: Ma, Ziping, et al.
Veröffentlicht: (2024)
von: Ma, Ziping, et al.
Veröffentlicht: (2024)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
Training-Free Reasoning and Reflection in MLLMs
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
von: Xiao, Kelaiti, et al.
Veröffentlicht: (2025)
Could Thinking Multilingually Empower LLM Reasoning?
von: Gao, Changjiang, et al.
Veröffentlicht: (2025)
von: Gao, Changjiang, et al.
Veröffentlicht: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
von: Zheng, Ruobing, et al.
Veröffentlicht: (2026) -
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
von: AI, Inclusion, et al.
Veröffentlicht: (2025) -
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025) -
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025) -
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025)