On Path to Multimodal Historical Reasoning: HistBench and HistAgent
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909653658501120 |
|---|---|
| author | Qiu, Jiahao Xiao, Fulian Wang, Yimin Mao, Yuchen Chen, Yijia Juan, Xinzhe Zhang, Shu Wang, Siran Qi, Xuan Zhang, Tongcheng Yao, Zixin Guo, Jiacheng Lu, Yifu Argon, Charles Cui, Jundi Chen, Daixin Zhou, Junran Zhou, Shuyao Zhou, Zhanpeng Yang, Ling Liu, Shilong Wang, Hongru Huang, Kaixuan Jiang, Xun Cao, Yuming Chen, Yue Chen, Yunfei Chen, Zhengyi Dai, Ruowei Deng, Mengqiu Fu, Jiye Gu, Yunting Guan, Zijie Huang, Zirui Ji, Xiaoyan Jiang, Yumeng Kong, Delong Li, Haolong Li, Jiaqi Li, Ruipeng Li, Tianze Li, Zhuoran Lian, Haixia Lin, Mengyue Liu, Xudong Lu, Jiayi Lu, Jinghan Luo, Wanyu Luo, Ziyue Pu, Zihao Qiao, Zhi Ren, Ruihuan Wan, Liang Wang, Ruixiang Wang, Tianhui Wang, Yang Wang, Zeyu Wang, Zihua Wu, Yujia Wu, Zhaoyi Xin, Hao Xing, Weiao Xiong, Ruojun Xu, Weijie Shu, Yao Xiao, Yao Yang, Xiaorui Yang, Yuchen Yi, Nan Yu, Jiadong Yu, Yangyuxuan Zeng, Huiting Zhang, Danni Zhang, Yunjie Zhang, Zhaoyu Zhang, Zhiheng Zheng, Xiaofeng Zhou, Peirong Zhong, Linyan Zong, Xiaoyin Zhao, Ying Chen, Zhenxin Ding, Lin Gao, Xiaoyu Gong, Bingbing Li, Yichao Liao, Yang Ma, Guang Ma, Tianyuan Sun, Xinrui Wang, Tianyi Xia, Han Xian, Ruobing Ye, Gen Yu, Tengfei Zhang, Wentao Wang, Yuxi Gao, Xi Wang, Mengdi |
| author_facet | Qiu, Jiahao Xiao, Fulian Wang, Yimin Mao, Yuchen Chen, Yijia Juan, Xinzhe Zhang, Shu Wang, Siran Qi, Xuan Zhang, Tongcheng Yao, Zixin Guo, Jiacheng Lu, Yifu Argon, Charles Cui, Jundi Chen, Daixin Zhou, Junran Zhou, Shuyao Zhou, Zhanpeng Yang, Ling Liu, Shilong Wang, Hongru Huang, Kaixuan Jiang, Xun Cao, Yuming Chen, Yue Chen, Yunfei Chen, Zhengyi Dai, Ruowei Deng, Mengqiu Fu, Jiye Gu, Yunting Guan, Zijie Huang, Zirui Ji, Xiaoyan Jiang, Yumeng Kong, Delong Li, Haolong Li, Jiaqi Li, Ruipeng Li, Tianze Li, Zhuoran Lian, Haixia Lin, Mengyue Liu, Xudong Lu, Jiayi Lu, Jinghan Luo, Wanyu Luo, Ziyue Pu, Zihao Qiao, Zhi Ren, Ruihuan Wan, Liang Wang, Ruixiang Wang, Tianhui Wang, Yang Wang, Zeyu Wang, Zihua Wu, Yujia Wu, Zhaoyi Xin, Hao Xing, Weiao Xiong, Ruojun Xu, Weijie Shu, Yao Xiao, Yao Yang, Xiaorui Yang, Yuchen Yi, Nan Yu, Jiadong Yu, Yangyuxuan Zeng, Huiting Zhang, Danni Zhang, Yunjie Zhang, Zhaoyu Zhang, Zhiheng Zheng, Xiaofeng Zhou, Peirong Zhong, Linyan Zong, Xiaoyin Zhao, Ying Chen, Zhenxin Ding, Lin Gao, Xiaoyu Gong, Bingbing Li, Yichao Liao, Yang Ma, Guang Ma, Tianyuan Sun, Xinrui Wang, Tianyi Xia, Han Xian, Ruobing Ye, Gen Yu, Tengfei Zhang, Wentao Wang, Yuxi Gao, Xi Wang, Mengdi |
| contents | Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents perform well on many existing benchmarks, they lack the domain-specific expertise required to engage with historical materials and questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality questions designed to evaluate AI's capacity for historical reasoning and authored by more than 40 expert contributors. The tasks span a wide range of historical problems-from factual retrieval based on primary sources to interpretive analysis of manuscripts and images, to interdisciplinary challenges involving archaeology, linguistics, or cultural history. Furthermore, the benchmark dataset spans 29 ancient and modern languages and covers a wide range of historical periods and world regions. Finding the poor performance of LLMs and other agents on HistBench, we further present HistAgent, a history-specific agent equipped with carefully designed tools for OCR, translation, archival search, and image understanding in History. On HistBench, HistAgent based on GPT-4o achieves an accuracy of 27.54% pass@1 and 36.47% pass@2, significantly outperforming LLMs with online search and generalist agents, including GPT-4o (18.60%), DeepSeek-R1(14.49%) and Open Deep Research-smolagents(20.29% pass@1 and 25.12% pass@2). These results highlight the limitations of existing LLMs and generalist agents and demonstrate the advantages of HistAgent for historical reasoning. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_20246 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | On Path to Multimodal Historical Reasoning: HistBench and HistAgent Qiu, Jiahao Xiao, Fulian Wang, Yimin Mao, Yuchen Chen, Yijia Juan, Xinzhe Zhang, Shu Wang, Siran Qi, Xuan Zhang, Tongcheng Yao, Zixin Guo, Jiacheng Lu, Yifu Argon, Charles Cui, Jundi Chen, Daixin Zhou, Junran Zhou, Shuyao Zhou, Zhanpeng Yang, Ling Liu, Shilong Wang, Hongru Huang, Kaixuan Jiang, Xun Cao, Yuming Chen, Yue Chen, Yunfei Chen, Zhengyi Dai, Ruowei Deng, Mengqiu Fu, Jiye Gu, Yunting Guan, Zijie Huang, Zirui Ji, Xiaoyan Jiang, Yumeng Kong, Delong Li, Haolong Li, Jiaqi Li, Ruipeng Li, Tianze Li, Zhuoran Lian, Haixia Lin, Mengyue Liu, Xudong Lu, Jiayi Lu, Jinghan Luo, Wanyu Luo, Ziyue Pu, Zihao Qiao, Zhi Ren, Ruihuan Wan, Liang Wang, Ruixiang Wang, Tianhui Wang, Yang Wang, Zeyu Wang, Zihua Wu, Yujia Wu, Zhaoyi Xin, Hao Xing, Weiao Xiong, Ruojun Xu, Weijie Shu, Yao Xiao, Yao Yang, Xiaorui Yang, Yuchen Yi, Nan Yu, Jiadong Yu, Yangyuxuan Zeng, Huiting Zhang, Danni Zhang, Yunjie Zhang, Zhaoyu Zhang, Zhiheng Zheng, Xiaofeng Zhou, Peirong Zhong, Linyan Zong, Xiaoyin Zhao, Ying Chen, Zhenxin Ding, Lin Gao, Xiaoyu Gong, Bingbing Li, Yichao Liao, Yang Ma, Guang Ma, Tianyuan Sun, Xinrui Wang, Tianyi Xia, Han Xian, Ruobing Ye, Gen Yu, Tengfei Zhang, Wentao Wang, Yuxi Gao, Xi Wang, Mengdi Artificial Intelligence Computation and Language Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents perform well on many existing benchmarks, they lack the domain-specific expertise required to engage with historical materials and questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality questions designed to evaluate AI's capacity for historical reasoning and authored by more than 40 expert contributors. The tasks span a wide range of historical problems-from factual retrieval based on primary sources to interpretive analysis of manuscripts and images, to interdisciplinary challenges involving archaeology, linguistics, or cultural history. Furthermore, the benchmark dataset spans 29 ancient and modern languages and covers a wide range of historical periods and world regions. Finding the poor performance of LLMs and other agents on HistBench, we further present HistAgent, a history-specific agent equipped with carefully designed tools for OCR, translation, archival search, and image understanding in History. On HistBench, HistAgent based on GPT-4o achieves an accuracy of 27.54% pass@1 and 36.47% pass@2, significantly outperforming LLMs with online search and generalist agents, including GPT-4o (18.60%), DeepSeek-R1(14.49%) and Open Deep Research-smolagents(20.29% pass@1 and 25.12% pass@2). These results highlight the limitations of existing LLMs and generalist agents and demonstrate the advantages of HistAgent for historical reasoning. |
| title | On Path to Multimodal Historical Reasoning: HistBench and HistAgent |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2505.20246 |