ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dai, Xinbang, Hu, Huikang, Chen, Yongrui, Li, Jiaqi, Jin, Rihui, Zhang, Yuyang, Li, Xiaoguang, Shang, Lifeng, Qi, Guilin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908750458126336
author Dai, Xinbang
Hu, Huikang
Chen, Yongrui
Li, Jiaqi
Jin, Rihui
Zhang, Yuyang
Li, Xiaoguang
Shang, Lifeng
Qi, Guilin
author_facet Dai, Xinbang
Hu, Huikang
Chen, Yongrui
Li, Jiaqi
Jin, Rihui
Zhang, Yuyang
Li, Xiaoguang
Shang, Lifeng
Qi, Guilin
contents While large language models (LLMs) excel at many domain-specific tasks, their ability to deeply comprehend and reason about full-length academic papers remains underexplored. Existing benchmarks often fall short of capturing such depth, either due to surface-level question design or unreliable evaluation metrics. To address this gap, we introduce ELAIPBench, a benchmark curated by domain experts to evaluate LLMs' comprehension of artificial intelligence (AI) research papers. Developed through an incentive-driven, adversarial annotation process, ELAIPBench features 403 multiple-choice questions from 137 papers. It spans three difficulty levels and emphasizes non-trivial reasoning rather than shallow retrieval. Our experiments show that the best-performing LLM achieves an accuracy of only 39.95%, far below human performance. Moreover, we observe that frontier LLMs equipped with a thinking mode or a retrieval-augmented generation (RAG) system fail to improve final results-even harming accuracy due to overthinking or noisy retrieval. These findings underscore the significant gap between current LLM capabilities and genuine comprehension of academic papers.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10549
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding
Dai, Xinbang
Hu, Huikang
Chen, Yongrui
Li, Jiaqi
Jin, Rihui
Zhang, Yuyang
Li, Xiaoguang
Shang, Lifeng
Qi, Guilin
Artificial Intelligence
While large language models (LLMs) excel at many domain-specific tasks, their ability to deeply comprehend and reason about full-length academic papers remains underexplored. Existing benchmarks often fall short of capturing such depth, either due to surface-level question design or unreliable evaluation metrics. To address this gap, we introduce ELAIPBench, a benchmark curated by domain experts to evaluate LLMs' comprehension of artificial intelligence (AI) research papers. Developed through an incentive-driven, adversarial annotation process, ELAIPBench features 403 multiple-choice questions from 137 papers. It spans three difficulty levels and emphasizes non-trivial reasoning rather than shallow retrieval. Our experiments show that the best-performing LLM achieves an accuracy of only 39.95%, far below human performance. Moreover, we observe that frontier LLMs equipped with a thinking mode or a retrieval-augmented generation (RAG) system fail to improve final results-even harming accuracy due to overthinking or noisy retrieval. These findings underscore the significant gap between current LLM capabilities and genuine comprehension of academic papers.
title ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding
topic Artificial Intelligence
url https://arxiv.org/abs/2510.10549