Saved in:
Bibliographic Details
Main Authors: Li, Annan, Wu, Chufan, Ge, Zengle, Chong, Yee Hin, Hou, Zhinan, Cao, Lizhe, Ju, Cheng, Wu, Jianmin, Li, Huaiming, Zhang, Haobo, Feng, Shenghao, Zhao, Mo, Qiu, Fengzhi, Yang, Rui, Zhang, Mengmeng, Zhu, Wenyi, Sun, Yingying, Sun, Quan, Yan, Shunhao, Liu, Danyu, Yin, Dawei, Shen, Dou
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.26144
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911474276892672
author Li, Annan
Wu, Chufan
Ge, Zengle
Chong, Yee Hin
Hou, Zhinan
Cao, Lizhe
Ju, Cheng
Wu, Jianmin
Li, Huaiming
Zhang, Haobo
Feng, Shenghao
Zhao, Mo
Qiu, Fengzhi
Yang, Rui
Zhang, Mengmeng
Zhu, Wenyi
Sun, Yingying
Sun, Quan
Yan, Shunhao
Liu, Danyu
Yin, Dawei
Shen, Dou
author_facet Li, Annan
Wu, Chufan
Ge, Zengle
Chong, Yee Hin
Hou, Zhinan
Cao, Lizhe
Ju, Cheng
Wu, Jianmin
Li, Huaiming
Zhang, Haobo
Feng, Shenghao
Zhao, Mo
Qiu, Fengzhi
Yang, Rui
Zhang, Mengmeng
Zhu, Wenyi
Sun, Yingying
Sun, Quan
Yan, Shunhao
Liu, Danyu
Yin, Dawei
Shen, Dou
contents Large language models (LLMs) are catalyzing the development of autonomous AI research agents for scientific and engineering discovery. We present FM Agent, a novel and general-purpose multi-agent framework that leverages a synergistic combination of LLM-based reasoning and large-scale evolutionary search to address complex real-world challenges. The core of FM Agent integrates several key innovations: 1) a cold-start initialization phase incorporating expert guidance, 2) a novel evolutionary sampling strategy for iterative optimization, 3) domain-specific evaluators that combine correctness, effectiveness, and LLM-supervised feedback, and 4) a distributed, asynchronous execution infrastructure built on Ray. Demonstrating broad applicability, our system has been evaluated across diverse domains, including operations research, machine learning, GPU kernel optimization, and classical mathematical problems. FM Agent reaches state-of-the-art results autonomously, without human interpretation or tuning -- 1976.3 on ALE-Bench (+5.2\%), 43.56\% on MLE-Bench (+4.0pp), up to 20x speedups on KernelBench, and establishes new state-of-the-art(SOTA) results on several classical mathematical problems. Beyond academic benchmarks, FM Agent shows considerable promise for both large-scale enterprise R\&D workflows and fundamental scientific research, where it can accelerate innovation, automate complex discovery processes, and deliver substantial engineering and scientific advances with broader societal impact.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The FM Agent
Li, Annan
Wu, Chufan
Ge, Zengle
Chong, Yee Hin
Hou, Zhinan
Cao, Lizhe
Ju, Cheng
Wu, Jianmin
Li, Huaiming
Zhang, Haobo
Feng, Shenghao
Zhao, Mo
Qiu, Fengzhi
Yang, Rui
Zhang, Mengmeng
Zhu, Wenyi
Sun, Yingying
Sun, Quan
Yan, Shunhao
Liu, Danyu
Yin, Dawei
Shen, Dou
Artificial Intelligence
Large language models (LLMs) are catalyzing the development of autonomous AI research agents for scientific and engineering discovery. We present FM Agent, a novel and general-purpose multi-agent framework that leverages a synergistic combination of LLM-based reasoning and large-scale evolutionary search to address complex real-world challenges. The core of FM Agent integrates several key innovations: 1) a cold-start initialization phase incorporating expert guidance, 2) a novel evolutionary sampling strategy for iterative optimization, 3) domain-specific evaluators that combine correctness, effectiveness, and LLM-supervised feedback, and 4) a distributed, asynchronous execution infrastructure built on Ray. Demonstrating broad applicability, our system has been evaluated across diverse domains, including operations research, machine learning, GPU kernel optimization, and classical mathematical problems. FM Agent reaches state-of-the-art results autonomously, without human interpretation or tuning -- 1976.3 on ALE-Bench (+5.2\%), 43.56\% on MLE-Bench (+4.0pp), up to 20x speedups on KernelBench, and establishes new state-of-the-art(SOTA) results on several classical mathematical problems. Beyond academic benchmarks, FM Agent shows considerable promise for both large-scale enterprise R\&D workflows and fundamental scientific research, where it can accelerate innovation, automate complex discovery processes, and deliver substantial engineering and scientific advances with broader societal impact.
title The FM Agent
topic Artificial Intelligence
url https://arxiv.org/abs/2510.26144