TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wei, Shaohang, Li, Wei, Song, Feifan, Luo, Wen, Zhuang, Tianyi, Tan, Haochen, Guo, Zhijiang, Wang, Houfeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918156097814528
author Wei, Shaohang
Li, Wei
Song, Feifan
Luo, Wen
Zhuang, Tianyi
Tan, Haochen
Guo, Zhijiang
Wang, Houfeng
author_facet Wei, Shaohang
Li, Wei
Song, Feifan
Luo, Wen
Zhuang, Tianyi
Tan, Haochen
Guo, Zhijiang
Wang, Houfeng
contents Temporal reasoning is pivotal for Large Language Models (LLMs) to comprehend the real world. However, existing works neglect the real-world challenges for temporal reasoning: (1) intensive temporal information, (2) fast-changing event dynamics, and (3) complex temporal dependencies in social interactions. To bridge this gap, we propose a multi-level benchmark TIME, designed for temporal reasoning in real-world scenarios. TIME consists of 38,522 QA pairs, covering 3 levels with 11 fine-grained sub-tasks. This benchmark encompasses 3 sub-datasets reflecting different real-world challenges: TIME-Wiki, TIME-News, and TIME-Dial. We conduct extensive experiments on reasoning models and non-reasoning models. And we conducted an in-depth analysis of temporal reasoning performance across diverse real-world scenarios and tasks, and summarized the impact of test-time scaling on temporal reasoning capabilities. Additionally, we release TIME-Lite, a human-annotated subset to foster future research and standardized evaluation in temporal reasoning. The code is available at https://github.com/sylvain-wei/TIME , the dataset is available at https://huggingface.co/datasets/SylvainWei/TIME , and the project page link is https://sylvain-wei.github.io/TIME/ .
format Preprint
id arxiv_https___arxiv_org_abs_2505_12891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
Wei, Shaohang
Li, Wei
Song, Feifan
Luo, Wen
Zhuang, Tianyi
Tan, Haochen
Guo, Zhijiang
Wang, Houfeng
Artificial Intelligence
Computation and Language
Temporal reasoning is pivotal for Large Language Models (LLMs) to comprehend the real world. However, existing works neglect the real-world challenges for temporal reasoning: (1) intensive temporal information, (2) fast-changing event dynamics, and (3) complex temporal dependencies in social interactions. To bridge this gap, we propose a multi-level benchmark TIME, designed for temporal reasoning in real-world scenarios. TIME consists of 38,522 QA pairs, covering 3 levels with 11 fine-grained sub-tasks. This benchmark encompasses 3 sub-datasets reflecting different real-world challenges: TIME-Wiki, TIME-News, and TIME-Dial. We conduct extensive experiments on reasoning models and non-reasoning models. And we conducted an in-depth analysis of temporal reasoning performance across diverse real-world scenarios and tasks, and summarized the impact of test-time scaling on temporal reasoning capabilities. Additionally, we release TIME-Lite, a human-annotated subset to foster future research and standardized evaluation in temporal reasoning. The code is available at https://github.com/sylvain-wei/TIME , the dataset is available at https://huggingface.co/datasets/SylvainWei/TIME , and the project page link is https://sylvain-wei.github.io/TIME/ .
title TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.12891