ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gui, Xin, Zhu, King, Ren, JinCheng, Chen, Qianben, Wang, Zekun Moore, LI, Yizhi, Liu, Xinpeng, Li, Xiaowan, Ren, Wenli, Miao, Linyu, Qin, Tianrui, Shu, Ziqi, Zhu, He, Tang, Xiangru, Shi, Dingfeng, Liu, Jiaheng, Jiang, Yuchen Eleanor, Liu, Minghao, Zhang, Ge, Zhou, Wangchunshu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914090627104768
author Gui, Xin
Zhu, King
Ren, JinCheng
Chen, Qianben
Wang, Zekun Moore
LI, Yizhi
Liu, Xinpeng
Li, Xiaowan
Ren, Wenli
Miao, Linyu
Qin, Tianrui
Shu, Ziqi
Zhu, He
Tang, Xiangru
Shi, Dingfeng
Liu, Jiaheng
Jiang, Yuchen Eleanor
Liu, Minghao
Zhang, Ge
Zhou, Wangchunshu
author_facet Gui, Xin
Zhu, King
Ren, JinCheng
Chen, Qianben
Wang, Zekun Moore
LI, Yizhi
Liu, Xinpeng
Li, Xiaowan
Ren, Wenli
Miao, Linyu
Qin, Tianrui
Shu, Ziqi
Zhu, He
Tang, Xiangru
Shi, Dingfeng
Liu, Jiaheng
Jiang, Yuchen Eleanor
Liu, Minghao
Zhang, Ge
Zhou, Wangchunshu
contents In recent years, the research focus of large language models (LLMs) and agents has shifted increasingly from demonstrating novel capabilities to complex reasoning and tackling challenging tasks. However, existing evaluations focus mainly on math/code contests or general tasks, while existing multi-domain academic benchmarks lack sufficient reasoning depth, leaving the field without a rigorous benchmark for high-level reasoning. To fill this gap, we introduce the Acadreason benchmark, designed to evaluate the ability of LLMs and agents to acquire and reason over academic knowledge. It consists of 50 expert-annotated academic problems across five high-reasoning domains, including computer science, economics, law, mathematics, and philosophy. All questions are sourced from top-tier publications in recent years and undergo rigorous annotation and quality control to ensure they are both challenging and answerable. We conduct systematic evaluations of over 10 mainstream LLMs and agents. The results show that most LLMs scored below 20 points, with even the cutting-edge GPT-5 achieving only 16 points. While agents achieved higher scores, none exceeded 40 points. This demonstrates the current capability gap between LLMs and agents in super-intelligent academic research tasks and highlights the challenges of Acadreason.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11652
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
Gui, Xin
Zhu, King
Ren, JinCheng
Chen, Qianben
Wang, Zekun Moore
LI, Yizhi
Liu, Xinpeng
Li, Xiaowan
Ren, Wenli
Miao, Linyu
Qin, Tianrui
Shu, Ziqi
Zhu, He
Tang, Xiangru
Shi, Dingfeng
Liu, Jiaheng
Jiang, Yuchen Eleanor
Liu, Minghao
Zhang, Ge
Zhou, Wangchunshu
Computation and Language
In recent years, the research focus of large language models (LLMs) and agents has shifted increasingly from demonstrating novel capabilities to complex reasoning and tackling challenging tasks. However, existing evaluations focus mainly on math/code contests or general tasks, while existing multi-domain academic benchmarks lack sufficient reasoning depth, leaving the field without a rigorous benchmark for high-level reasoning. To fill this gap, we introduce the Acadreason benchmark, designed to evaluate the ability of LLMs and agents to acquire and reason over academic knowledge. It consists of 50 expert-annotated academic problems across five high-reasoning domains, including computer science, economics, law, mathematics, and philosophy. All questions are sourced from top-tier publications in recent years and undergo rigorous annotation and quality control to ensure they are both challenging and answerable. We conduct systematic evaluations of over 10 mainstream LLMs and agents. The results show that most LLMs scored below 20 points, with even the cutting-edge GPT-5 achieving only 16 points. While agents achieved higher scores, none exceeded 40 points. This demonstrates the current capability gap between LLMs and agents in super-intelligent academic research tasks and highlights the challenges of Acadreason.
title ACADREASON: Exploring the Limits of Reasoning Models with Academic Research Problems
topic Computation and Language
url https://arxiv.org/abs/2510.11652