_version_ 1866911612385886208
author MiroMind Team
Bai, Song
Bing, Lidong
Chen, Carson
Chen, Guanzheng
Chen, Yuntao
Chen, Zhe
Chen, Ziyi
Dai, Jifeng
Dong, Xuan
Dou, Wenhan
Deng, Yue
Fu, Yunjie
Ge, Junqi
Han, Chenxia
Huang, Tammy
Huang, Zhenhang
Jiao, Jerry
Jiang, Shilei
Jiao, Tianyu
Jian, Xiaoqi
Lei, Lei
Li, Ruilin
Luo, Gen
Li, Tiantong
Lin, Xiang
Liu, Ziyuan
Li, Zhiqi
Ni, Jie
Ren, Qiang
Sun, Pax
Su, Shiqian
Tao, Chenxin
Wang, Bin
Wang, Wenhai
Wang, Haonan
Wang, James
Wang, Jin
Wang, Jojo
Wang, Letian
Wang, Shizun
Wang, Weizhi
Wang, Zixuan
Xu, Jinfan
Xing, Sen
Yang, Chenyu
Ye, Hai
Yu, Jiaheng
Yu, Yue
Zhong, Muyan
Zhao, Tianchen
Zhu, Xizhou
Zhou, Yanpeng
Zhang, Yifan
Zhu, Zhi
author_facet MiroMind Team
Bai, Song
Bing, Lidong
Chen, Carson
Chen, Guanzheng
Chen, Yuntao
Chen, Zhe
Chen, Ziyi
Dai, Jifeng
Dong, Xuan
Dou, Wenhan
Deng, Yue
Fu, Yunjie
Ge, Junqi
Han, Chenxia
Huang, Tammy
Huang, Zhenhang
Jiao, Jerry
Jiang, Shilei
Jiao, Tianyu
Jian, Xiaoqi
Lei, Lei
Li, Ruilin
Luo, Gen
Li, Tiantong
Lin, Xiang
Liu, Ziyuan
Li, Zhiqi
Ni, Jie
Ren, Qiang
Sun, Pax
Su, Shiqian
Tao, Chenxin
Wang, Bin
Wang, Wenhai
Wang, Haonan
Wang, James
Wang, Jin
Wang, Jojo
Wang, Letian
Wang, Shizun
Wang, Weizhi
Wang, Zixuan
Xu, Jinfan
Xing, Sen
Yang, Chenyu
Ye, Hai
Yu, Jiaheng
Yu, Yue
Zhong, Muyan
Zhao, Tianchen
Zhu, Xizhou
Zhou, Yanpeng
Zhang, Yifan
Zhu, Zhi
contents We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale up model size or context length, MiroThinker explores interaction scaling at the model level, systematically training the model to handle deeper and more frequent agent-environment interactions as a third dimension of performance improvement. Unlike LLM test-time scaling, which operates in isolation and risks degradation with longer reasoning chains, interactive scaling leverages environment feedback and external information acquisition to correct errors and refine trajectories. Through reinforcement learning, the model achieves efficient interaction scaling: with a 256K context window, it can perform up to 600 tool calls per task, enabling sustained multi-turn reasoning and complex real-world research workflows. Across four representative benchmarks-GAIA, HLE, BrowseComp, and BrowseComp-ZH-the 72B variant achieves up to 81.9%, 37.7%, 47.1%, and 55.6% accuracy respectively, surpassing previous open-source agents and approaching commercial counterparts such as GPT-5-high. Our analysis reveals that MiroThinker benefits from interactive scaling consistently: research performance improves predictably as the model engages in deeper and more frequent agent-environment interactions, demonstrating that interaction depth exhibits scaling behaviors analogous to model size and context length. These findings establish interaction scaling as a third critical dimension for building next-generation open research agents, complementing model capacity and context windows.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
MiroMind Team
Bai, Song
Bing, Lidong
Chen, Carson
Chen, Guanzheng
Chen, Yuntao
Chen, Zhe
Chen, Ziyi
Dai, Jifeng
Dong, Xuan
Dou, Wenhan
Deng, Yue
Fu, Yunjie
Ge, Junqi
Han, Chenxia
Huang, Tammy
Huang, Zhenhang
Jiao, Jerry
Jiang, Shilei
Jiao, Tianyu
Jian, Xiaoqi
Lei, Lei
Li, Ruilin
Luo, Gen
Li, Tiantong
Lin, Xiang
Liu, Ziyuan
Li, Zhiqi
Ni, Jie
Ren, Qiang
Sun, Pax
Su, Shiqian
Tao, Chenxin
Wang, Bin
Wang, Wenhai
Wang, Haonan
Wang, James
Wang, Jin
Wang, Jojo
Wang, Letian
Wang, Shizun
Wang, Weizhi
Wang, Zixuan
Xu, Jinfan
Xing, Sen
Yang, Chenyu
Ye, Hai
Yu, Jiaheng
Yu, Yue
Zhong, Muyan
Zhao, Tianchen
Zhu, Xizhou
Zhou, Yanpeng
Zhang, Yifan
Zhu, Zhi
Computation and Language
We present MiroThinker v1.0, an open-source research agent designed to advance tool-augmented reasoning and information-seeking capabilities. Unlike previous agents that only scale up model size or context length, MiroThinker explores interaction scaling at the model level, systematically training the model to handle deeper and more frequent agent-environment interactions as a third dimension of performance improvement. Unlike LLM test-time scaling, which operates in isolation and risks degradation with longer reasoning chains, interactive scaling leverages environment feedback and external information acquisition to correct errors and refine trajectories. Through reinforcement learning, the model achieves efficient interaction scaling: with a 256K context window, it can perform up to 600 tool calls per task, enabling sustained multi-turn reasoning and complex real-world research workflows. Across four representative benchmarks-GAIA, HLE, BrowseComp, and BrowseComp-ZH-the 72B variant achieves up to 81.9%, 37.7%, 47.1%, and 55.6% accuracy respectively, surpassing previous open-source agents and approaching commercial counterparts such as GPT-5-high. Our analysis reveals that MiroThinker benefits from interactive scaling consistently: research performance improves predictably as the model engages in deeper and more frequent agent-environment interactions, demonstrating that interaction depth exhibits scaling behaviors analogous to model size and context length. These findings establish interaction scaling as a third critical dimension for building next-generation open research agents, complementing model capacity and context windows.
title MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
topic Computation and Language
url https://arxiv.org/abs/2511.11793