LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Weizhi, Nuamah, Kwabena, Belle, Vaishak
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914233453641728
author Tang, Weizhi
Nuamah, Kwabena
Belle, Vaishak
author_facet Tang, Weizhi
Nuamah, Kwabena
Belle, Vaishak
contents Temporal Reasoning (TR) is a critical ability for LLMs to understand and reason over temporal information and relationships between events. To study the TR ability in LLMs, prior works provide different ways for evaluating various aspects of TR ability. In this work, we propose an alternative perspective for evaluating TR ability by leveraging Linear Temporal Logic (LTL), and develop a pipeline to automatically synthesize challenges for assessing the TR ability of LLMs. Based on this pipeline, we construct a dataset, namely LTLBench, consisting of $2000$ TR challenges, and benchmark 12 LLMs across 5 different methods. Furthermore, we conduct additional experiments to investigate the impact of increasing the number of formula operators and events on both LLM performance and the complexity of TR problems. We also perform qualitative analyses of their reasoning processes and the effects of varying the number of events and formula operators, which reveal 3 main issues in their temporal reasoning processes and the unexpected performance changes observed as problem complexity increases. We expect this work to provide valuable insights into the TR ability of LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05434
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
Tang, Weizhi
Nuamah, Kwabena
Belle, Vaishak
Computation and Language
Artificial Intelligence
Temporal Reasoning (TR) is a critical ability for LLMs to understand and reason over temporal information and relationships between events. To study the TR ability in LLMs, prior works provide different ways for evaluating various aspects of TR ability. In this work, we propose an alternative perspective for evaluating TR ability by leveraging Linear Temporal Logic (LTL), and develop a pipeline to automatically synthesize challenges for assessing the TR ability of LLMs. Based on this pipeline, we construct a dataset, namely LTLBench, consisting of $2000$ TR challenges, and benchmark 12 LLMs across 5 different methods. Furthermore, we conduct additional experiments to investigate the impact of increasing the number of formula operators and events on both LLM performance and the complexity of TR problems. We also perform qualitative analyses of their reasoning processes and the effects of varying the number of events and formula operators, which reveal 3 main issues in their temporal reasoning processes and the unexpected performance changes observed as problem complexity increases. We expect this work to provide valuable insights into the TR ability of LLMs.
title LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.05434