TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Zheng, Chen, Jingchang, Chen, Qianglong, Yu, Weijiang, Wang, Haotian, Liu, Ming, Qin, Bing
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909232667820032
author Chu, Zheng
Chen, Jingchang
Chen, Qianglong
Yu, Weijiang
Wang, Haotian
Liu, Ming
Qin, Bing
author_facet Chu, Zheng
Chen, Jingchang
Chen, Qianglong
Yu, Weijiang
Wang, Haotian
Liu, Ming
Qin, Bing
contents Grasping the concept of time is a fundamental facet of human cognition, indispensable for truly comprehending the intricacies of the world. Previous studies typically focus on specific aspects of time, lacking a comprehensive temporal reasoning benchmark. To address this, we propose TimeBench, a comprehensive hierarchical temporal reasoning benchmark that covers a broad spectrum of temporal reasoning phenomena. TimeBench provides a thorough evaluation for investigating the temporal reasoning capabilities of large language models. We conduct extensive experiments on GPT-4, LLaMA2, and other popular LLMs under various settings. Our experimental results indicate a significant performance gap between the state-of-the-art LLMs and humans, highlighting that there is still a considerable distance to cover in temporal reasoning. Besides, LLMs exhibit capability discrepancies across different reasoning categories. Furthermore, we thoroughly analyze the impact of multiple aspects on temporal reasoning and emphasize the associated challenges. We aspire for TimeBench to serve as a comprehensive benchmark, fostering research in temporal reasoning. Resources are available at: https://github.com/zchuz/TimeBench
format Preprint
id arxiv_https___arxiv_org_abs_2311_17667
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
Chu, Zheng
Chen, Jingchang
Chen, Qianglong
Yu, Weijiang
Wang, Haotian
Liu, Ming
Qin, Bing
Computation and Language
Artificial Intelligence
Grasping the concept of time is a fundamental facet of human cognition, indispensable for truly comprehending the intricacies of the world. Previous studies typically focus on specific aspects of time, lacking a comprehensive temporal reasoning benchmark. To address this, we propose TimeBench, a comprehensive hierarchical temporal reasoning benchmark that covers a broad spectrum of temporal reasoning phenomena. TimeBench provides a thorough evaluation for investigating the temporal reasoning capabilities of large language models. We conduct extensive experiments on GPT-4, LLaMA2, and other popular LLMs under various settings. Our experimental results indicate a significant performance gap between the state-of-the-art LLMs and humans, highlighting that there is still a considerable distance to cover in temporal reasoning. Besides, LLMs exhibit capability discrepancies across different reasoning categories. Furthermore, we thoroughly analyze the impact of multiple aspects on temporal reasoning and emphasize the associated challenges. We aspire for TimeBench to serve as a comprehensive benchmark, fostering research in temporal reasoning. Resources are available at: https://github.com/zchuz/TimeBench
title TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.17667