Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gao, Lirong, Wang, Zeqing, Cai, Yuyan, Deng, Jiayi, Gu, Yanmei, Zhang, Yiming, Zhou, Jia, Zhang, Yanfei, Zhao, Junbo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918470164152320
author Gao, Lirong
Wang, Zeqing
Cai, Yuyan
Deng, Jiayi
Gu, Yanmei
Zhang, Yiming
Zhou, Jia
Zhang, Yanfei
Zhao, Junbo
author_facet Gao, Lirong
Wang, Zeqing
Cai, Yuyan
Deng, Jiayi
Gu, Yanmei
Zhang, Yiming
Zhou, Jia
Zhang, Yanfei
Zhao, Junbo
contents While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level historical reasoning remains underexplored. Existing benchmarks primarily assess basic knowledge breadth or lexical understanding, failing to capture the higher-order skills, such as evidentiary reasoning,that are central to historical research. To fill this gap, we introduce ProHist-Bench, a novel benchmark anchored in the Chinese Imperial Examination (Keju) system, a comprehensive microcosm of East Asian political, social, and intellectual history spanning over 1,300 years. Developed through deep interdisciplinary collaboration, ProHist-Bench features 400 challenging, expert-curated questions across eight dynasties, accompanied by 10,891 fine-grained evaluation rubrics. Through a rigorous evaluation of 18 LLMs, we reveal a significant proficiency gap: even state-of-the-art LLMs struggle with complex historical research questions. We hope ProHist-Bench will facilitate the development of domain-specific reasoning LLMs, advance computational historical research, and further uncover the untapped potential of LLMs. We release ProHist-Bench at https://github.com/inclusionAI/ABench/tree/main/ProHist-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24690
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
Gao, Lirong
Wang, Zeqing
Cai, Yuyan
Deng, Jiayi
Gu, Yanmei
Zhang, Yiming
Zhou, Jia
Zhang, Yanfei
Zhao, Junbo
Computation and Language
While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level historical reasoning remains underexplored. Existing benchmarks primarily assess basic knowledge breadth or lexical understanding, failing to capture the higher-order skills, such as evidentiary reasoning,that are central to historical research. To fill this gap, we introduce ProHist-Bench, a novel benchmark anchored in the Chinese Imperial Examination (Keju) system, a comprehensive microcosm of East Asian political, social, and intellectual history spanning over 1,300 years. Developed through deep interdisciplinary collaboration, ProHist-Bench features 400 challenging, expert-curated questions across eight dynasties, accompanied by 10,891 fine-grained evaluation rubrics. Through a rigorous evaluation of 18 LLMs, we reveal a significant proficiency gap: even state-of-the-art LLMs struggle with complex historical research questions. We hope ProHist-Bench will facilitate the development of domain-specific reasoning LLMs, advance computational historical research, and further uncover the untapped potential of LLMs. We release ProHist-Bench at https://github.com/inclusionAI/ABench/tree/main/ProHist-Bench.
title Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
topic Computation and Language
url https://arxiv.org/abs/2604.24690