SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
Fuente:
arXiv
Saved in:
| Main Authors: | Chai, Jingyi, Tang, Shuo, Ye, Rui, Du, Yuwen, Zhu, Xinyu, Zhou, Mengcheng, Wang, Yanfeng, E, Weinan, Zhang, Yuzhi, Zhang, Linfeng, Chen, Siheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale
by: Zhang, Linfeng, et al.
Published: (2025)
by: Zhang, Linfeng, et al.
Published: (2025)
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
by: Pang, Xianghe, et al.
Published: (2025)
by: Pang, Xianghe, et al.
Published: (2025)
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
by: Miao, Tingjia, et al.
Published: (2025)
by: Miao, Tingjia, et al.
Published: (2025)
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering
by: Zhu, Xinyu, et al.
Published: (2026)
by: Zhu, Xinyu, et al.
Published: (2026)
Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Incentivizing Inclusive Contributions in Model Sharing Markets
by: Zhang, Enpei, et al.
Published: (2025)
by: Zhang, Enpei, et al.
Published: (2025)
FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
DataMaster: Data-Centric Autonomous AI Research
by: Du, Yaxin, et al.
Published: (2026)
by: Du, Yaxin, et al.
Published: (2026)
Towards Self-Evolving Agentic Literature Retrieval
by: Du, Yuwen, et al.
Published: (2026)
by: Du, Yuwen, et al.
Published: (2026)
TeachMaster: Generative Teaching via Code
by: Wang, Yuheng, et al.
Published: (2025)
by: Wang, Yuheng, et al.
Published: (2025)
Emerging Safety Attack and Defense in Federated Instruction Tuning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
KnowledgeSG: Privacy-Preserving Synthetic Text Generation with Knowledge Distillation from Server
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Leveraging Unstructured Text Data for Federated Instruction Tuning of Large Language Models
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
by: Du, Yuwen, et al.
Published: (2026)
by: Du, Yuwen, et al.
Published: (2026)
OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories
by: Du, Yuwen, et al.
Published: (2026)
by: Du, Yuwen, et al.
Published: (2026)
Humanity's Last Exam
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
by: Liu, Yuzhi, et al.
Published: (2026)
by: Liu, Yuzhi, et al.
Published: (2026)
Enhancing Data Quality in Federated Fine-Tuning of Foundation Models
by: Zhao, Wanru, et al.
Published: (2024)
by: Zhao, Wanru, et al.
Published: (2024)
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?
by: Li, Xiangyang, et al.
Published: (2025)
by: Li, Xiangyang, et al.
Published: (2025)
Can AI Master Construction Management (CM)? Benchmarking State-of-the-Art Large Language Models on CM Certification Exams
by: Xiong, Ruoxin, et al.
Published: (2025)
by: Xiong, Ruoxin, et al.
Published: (2025)
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
AGI-Elo: How Far Are We From Mastering A Task?
by: Sun, Shuo, et al.
Published: (2025)
by: Sun, Shuo, et al.
Published: (2025)
Ep. 119: The 1% Rule: Mastering Kaizen for Lasting Improvement
by: Rosehill, Daniel, et al.
Published: (2025)
by: Rosehill, Daniel, et al.
Published: (2025)
Have We Mastered Scale in Deep Monocular Visual SLAM? The ScaleMaster Dataset and Benchmark
by: Ju, Hyoseok, et al.
Published: (2026)
by: Ju, Hyoseok, et al.
Published: (2026)
MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
by: Wang, Wenhao, et al.
Published: (2026)
by: Wang, Wenhao, et al.
Published: (2026)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
by: Ye, Rui, et al.
Published: (2024)
by: Ye, Rui, et al.
Published: (2024)
Grammar Module - Mastering the Parts of Speech
by: ARPON, EMELY
Published: (2025)
by: ARPON, EMELY
Published: (2025)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
by: Li, Sihang, et al.
Published: (2024)
by: Li, Sihang, et al.
Published: (2024)
HLL: Can Agents Cross Humanity's Last Line of Verification?
by: Song, Xinhao, et al.
Published: (2026)
by: Song, Xinhao, et al.
Published: (2026)
Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
by: Gallouédec, Quentin, et al.
Published: (2024)
by: Gallouédec, Quentin, et al.
Published: (2024)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
by: Guan, Xinyu, et al.
Published: (2025)
by: Guan, Xinyu, et al.
Published: (2025)
Uni-Mol3: A Multi-Molecular Foundation Model for Advancing Organic Reaction Modeling
by: Wu, Lirong, et al.
Published: (2025)
by: Wu, Lirong, et al.
Published: (2025)
Can Large Language Models Master Complex Card Games?
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
On the Last Kervaire Invariant Problem
by: Lin, Weinan, et al.
Published: (2024)
by: Lin, Weinan, et al.
Published: (2024)
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
Similar Items
-
Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale
by: Zhang, Linfeng, et al.
Published: (2025) -
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
by: Zhu, Xinyu, et al.
Published: (2026) -
BrowseMaster: Towards Scalable Web Browsing via Tool-Augmented Programmatic Agent Pair
by: Pang, Xianghe, et al.
Published: (2025) -
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
by: Miao, Tingjia, et al.
Published: (2025) -
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
by: Liu, Zexi, et al.
Published: (2025)