Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Lirong, Wang, Zeqing, Cai, Yuyan, Deng, Jiayi, Gu, Yanmei, Zhang, Yiming, Zhou, Jia, Zhang, Yanfei, Zhao, Junbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
von: Dong, Peijie, et al.
Veröffentlicht: (2025)
von: Dong, Peijie, et al.
Veröffentlicht: (2025)
DORY: Deliberative Prompt Recovery for LLM
von: Gao, Lirong, et al.
Veröffentlicht: (2024)
von: Gao, Lirong, et al.
Veröffentlicht: (2024)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
von: Zhou, Ziya, et al.
Veröffentlicht: (2024)
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Evaluating LLMs on Chinese Idiom Translation
von: Yang, Cai, et al.
Veröffentlicht: (2025)
von: Yang, Cai, et al.
Veröffentlicht: (2025)
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
A Thorough Examination of Decoding Methods in the Era of LLMs
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
von: Zhang, Yiming, et al.
Veröffentlicht: (2025)
Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language
von: Lin, Dianqing, et al.
Veröffentlicht: (2026)
von: Lin, Dianqing, et al.
Veröffentlicht: (2026)
3D-PreMise: Can Large Language Models Generate 3D Shapes with Sharp Features and Parametric Control?
von: Yuan, Zeqing, et al.
Veröffentlicht: (2024)
von: Yuan, Zeqing, et al.
Veröffentlicht: (2024)
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
von: Miao, Tingjia, et al.
Veröffentlicht: (2026)
von: Miao, Tingjia, et al.
Veröffentlicht: (2026)
When LLMs Can't Help: Real-World Evaluation of LLMs in Nutrition
von: Li, Karen Jia-Hui, et al.
Veröffentlicht: (2025)
von: Li, Karen Jia-Hui, et al.
Veröffentlicht: (2025)
How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?
von: Zheng, Qiaoyu, et al.
Veröffentlicht: (2024)
von: Zheng, Qiaoyu, et al.
Veröffentlicht: (2024)
Collaborative QA using Interacting LLMs. Impact of Network Structure, Node Capability and Distributed Data
von: Jain, Adit, et al.
Veröffentlicht: (2025)
von: Jain, Adit, et al.
Veröffentlicht: (2025)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
von: Li, Chloe, et al.
Veröffentlicht: (2025)
von: Li, Chloe, et al.
Veröffentlicht: (2025)
Can LLMs Classify CVEs? Investigating LLMs Capabilities in Computing CVSS Vectors
von: Marchiori, Francesco, et al.
Veröffentlicht: (2025)
von: Marchiori, Francesco, et al.
Veröffentlicht: (2025)
Enough Coin Flips Can Make LLMs Act Bayesian
von: Gupta, Ritwik, et al.
Veröffentlicht: (2025)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2025)
Towards Automatic Evaluation for LLMs' Clinical Capabilities: Metric, Data, and Algorithm
von: Liu, Lei, et al.
Veröffentlicht: (2024)
von: Liu, Lei, et al.
Veröffentlicht: (2024)
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
von: Liu, Chuang, et al.
Veröffentlicht: (2024)
Evaluating Developmental Cognition Capabilities of LLMs
von: Xiao, Xiao, et al.
Veröffentlicht: (2026)
von: Xiao, Xiao, et al.
Veröffentlicht: (2026)
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
von: Zhu, Shaojie, et al.
Veröffentlicht: (2023)
von: Zhu, Shaojie, et al.
Veröffentlicht: (2023)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
von: Li, Haoming, et al.
Veröffentlicht: (2025)
von: Li, Haoming, et al.
Veröffentlicht: (2025)
Can LLMs be Fooled? Investigating Vulnerabilities in LLMs
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
von: Abdali, Sara, et al.
Veröffentlicht: (2024)
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
von: Zhang, Leizhen, et al.
Veröffentlicht: (2026)
von: Zhang, Leizhen, et al.
Veröffentlicht: (2026)
Typestate via Revocable Capabilities
von: Jia, Songlin, et al.
Veröffentlicht: (2025)
von: Jia, Songlin, et al.
Veröffentlicht: (2025)
An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
von: Chen, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Chen, Zhiyuan, et al.
Veröffentlicht: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
von: Tie, Guiyao, et al.
Veröffentlicht: (2025)
Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
von: Chen, Zejian, et al.
Veröffentlicht: (2026)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
von: Kwon, Deuksin, et al.
Veröffentlicht: (2024)
von: Kwon, Deuksin, et al.
Veröffentlicht: (2024)
Can Editing LLMs Inject Harm?
von: Chen, Canyu, et al.
Veröffentlicht: (2024)
von: Chen, Canyu, et al.
Veröffentlicht: (2024)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
von: Gao, Xin, et al.
Veröffentlicht: (2025)
von: Gao, Xin, et al.
Veröffentlicht: (2025)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Yao, Zhiyuan, et al.
Veröffentlicht: (2026)
Learners as Historians: Making History Come Alive through Historical Inquiry
von: Pappas, Marjorie L.
Veröffentlicht: (2007)
von: Pappas, Marjorie L.
Veröffentlicht: (2007)
SCAN: Structured Capability Assessment and Navigation for LLMs
von: Wang, Zongqi, et al.
Veröffentlicht: (2025)
von: Wang, Zongqi, et al.
Veröffentlicht: (2025)
LLMAID: Identifying AI Capabilities in Android Apps with LLMs
von: Liu, Pei, et al.
Veröffentlicht: (2025)
von: Liu, Pei, et al.
Veröffentlicht: (2025)
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
von: Wu, Yuyang, et al.
Veröffentlicht: (2026)
von: Wu, Yuyang, et al.
Veröffentlicht: (2026)
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
von: Son, Guijin, et al.
Veröffentlicht: (2026)
von: Son, Guijin, et al.
Veröffentlicht: (2026)
D.Va: Validate Your Demonstration First Before You Use It
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
von: Zhang, Qi, et al.
Veröffentlicht: (2025)
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
von: Guo, Zichun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
von: Dong, Peijie, et al.
Veröffentlicht: (2025) -
DORY: Deliberative Prompt Recovery for LLM
von: Gao, Lirong, et al.
Veröffentlicht: (2024) -
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
von: Zhou, Ziya, et al.
Veröffentlicht: (2024) -
FLoE: Fisher-Based Layer Selection for Efficient Sparse Adaptation of Low-Rank Experts
von: Wang, Xinyi, et al.
Veröffentlicht: (2025) -
Evaluating LLMs on Chinese Idiom Translation
von: Yang, Cai, et al.
Veröffentlicht: (2025)