MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xingxuan, Xiao, Yao, Ng, Dianwen, Ye, Hai, Deng, Yue, Lin, Xiang, Wang, Bin, Mo, Zhanfeng, Zhang, Chong, Zhang, Yueyi, Yang, Zonglin, Li, Ruilin, Lei, Lei, Xu, Shihao, Zhao, Han, Chen, Weiling, Ji, Feng, Bing, Lidong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Agent Tool-Integrated Policy Optimization
by: Mo, Zhanfeng, et al.
Published: (2025)
by: Mo, Zhanfeng, et al.
Published: (2025)
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
by: Zhang, Chong, et al.
Published: (2025)
by: Zhang, Chong, et al.
Published: (2025)
ParaICL: Towards Parallel In-Context Learning
by: Li, Xingxuan, et al.
Published: (2024)
by: Li, Xingxuan, et al.
Published: (2024)
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)
by: Li, Xingxuan, et al.
Published: (2022)
Unlocking Temporal Question Answering for Large Language Models with Tailor-Made Reasoning Logic
by: Li, Xingxuan, et al.
Published: (2023)
by: Li, Xingxuan, et al.
Published: (2023)
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
by: MiroMind Team, et al.
Published: (2025)
by: MiroMind Team, et al.
Published: (2025)
Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources
by: Li, Xingxuan, et al.
Published: (2023)
by: Li, Xingxuan, et al.
Published: (2023)
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
by: Li, Xingxuan, et al.
Published: (2024)
by: Li, Xingxuan, et al.
Published: (2024)
First Try Matters: Revisiting the Role of Reflection in Reasoning Models
by: Kang, Liwei, et al.
Published: (2025)
by: Kang, Liwei, et al.
Published: (2025)
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
by: Ye, Fangda, et al.
Published: (2026)
by: Ye, Fangda, et al.
Published: (2026)
MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks
by: Su, Shiqian, et al.
Published: (2026)
by: Su, Shiqian, et al.
Published: (2026)
Principal Context-aware Diffusion Guided Data Augmentation for Fault Localization
by: Fu, Shihao, et al.
Published: (2025)
by: Fu, Shihao, et al.
Published: (2025)
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
by: Chen, Guanzheng, et al.
Published: (2026)
by: Chen, Guanzheng, et al.
Published: (2026)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
Porous AgPd Nanomushrooms with Enhanced Methanol Tolerance for Oxygen Reduction Reaction
by: Yueyi Cui, et al.
Published: (2024)
by: Yueyi Cui, et al.
Published: (2024)
Towards Robust Temporal Reasoning of Large Language Models via a Multi-Hop QA Dataset and Pseudo-Instruction Tuning
by: Tan, Qingyu, et al.
Published: (2023)
by: Tan, Qingyu, et al.
Published: (2023)
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Towards a bottom-up formulation of spin kinetic theory
by: Mo, Zonglin, et al.
Published: (2025)
by: Mo, Zonglin, et al.
Published: (2025)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Landslide Inventory Mapping in the Upper Minjiang River, Eastern Tibetan Plateau, Using Multi‐Source Satellite Remote Sensing
by: Chong Geng, et al.
Published: (2025)
by: Chong Geng, et al.
Published: (2025)
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
TFMLinker: Universal Link Predictor by Graph In-Context Learning with Tabular Foundation Models
by: Liao, Tianyin, et al.
Published: (2026)
by: Liao, Tianyin, et al.
Published: (2026)
Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration
by: Liu, Jiaheng, et al.
Published: (2026)
by: Liu, Jiaheng, et al.
Published: (2026)
MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback
by: Liu, Wanhao, et al.
Published: (2025)
by: Liu, Wanhao, et al.
Published: (2025)
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
by: Zhou, Kun, et al.
Published: (2024)
by: Zhou, Kun, et al.
Published: (2024)
Secure Beamforming and Reflection Design for RIS-ISAC Systems Under Collusion of Passive and Active Eavesdroppers
by: Dong, Yueyi, et al.
Published: (2026)
by: Dong, Yueyi, et al.
Published: (2026)
Document Reconstruction Unlocks Scalable Long-Context RLVR
by: Xiao, Yao, et al.
Published: (2026)
by: Xiao, Yao, et al.
Published: (2026)
A Bounded Rationality Model for Sustainable Supplier Selection for New Products: A Whole Product Life Cycle Perspective
by: Chong Wu, et al.
Published: (2025)
by: Chong Wu, et al.
Published: (2025)
Oda a Joan Miró : litografies de Joan Miró / Joan Brossa
by: Brossa, Joan
Published: (1973)
by: Brossa, Joan
Published: (1973)
Biopolymeric Gels: Advancements in Sustainable Multifunctional Materials
by: Chuxin Lei, et al.
Published: (2025)
by: Chuxin Lei, et al.
Published: (2025)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
by: Shi, Wenhao, et al.
Published: (2024)
by: Shi, Wenhao, et al.
Published: (2024)
Fast Graph Generation via Spectral Diffusion
by: Luo, Tianze, et al.
Published: (2022)
by: Luo, Tianze, et al.
Published: (2022)
Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
Purification of mesenchymal stromal cell‐derived small extracellular vesicles using ultrafiltration
by: Rui Lei, et al.
Published: (2025)
by: Rui Lei, et al.
Published: (2025)
Room‐Temperature Degradation of Lignin from Wasted Seed Coats for the Production of Pyrocatechol/Gallol Derivatives
by: Shihao Su, et al.
Published: (2026)
by: Shihao Su, et al.
Published: (2026)
BEATS: An Open-Source, High-Precision, Multi-Channel EEG Acquisition Tool System
by: Zou, Bing, et al.
Published: (2022)
by: Zou, Bing, et al.
Published: (2022)
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification
by: MiroMind Team, et al.
Published: (2026)
by: MiroMind Team, et al.
Published: (2026)
Similar Items
-
Multi-Agent Tool-Integrated Policy Optimization
by: Mo, Zhanfeng, et al.
Published: (2025) -
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
by: Zhang, Chong, et al.
Published: (2025) -
ParaICL: Towards Parallel In-Context Learning
by: Li, Xingxuan, et al.
Published: (2024) -
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026) -
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)