TCP: a Benchmark for Temporal Constraint-Based Planning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ding, Zifeng, Yan, Sikuan, Yuan, Zhangdie, Hu, Xianglong, Lin, Fangru, Vlachos, Andreas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915547777597440
author Ding, Zifeng
Yan, Sikuan
Yuan, Zhangdie
Hu, Xianglong
Lin, Fangru
Vlachos, Andreas
author_facet Ding, Zifeng
Yan, Sikuan
Yuan, Zhangdie
Hu, Xianglong
Lin, Fangru
Vlachos, Andreas
contents Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity. To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark that jointly assesses both capabilities. Each instance in TCP features a naturalistic dialogue around a collaborative project, where diverse and interdependent temporal constraints are explicitly or implicitly expressed, and models must infer an optimal schedule that satisfies all constraints. To construct TCP, we generate abstract problem prototypes that are then paired with realistic scenarios from various domains and enriched into dialogues using an LLM. A human quality check is performed on a sampled subset to confirm the reliability of our benchmark. We evaluate state-of-the-art LLMs and find that even the strongest models may struggle with TCP, highlighting its difficulty and revealing limitations in LLMs' temporal constraint-based planning abilities. We analyze underlying failure cases, open source our benchmark, and hope our findings can inspire future research.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TCP: a Benchmark for Temporal Constraint-Based Planning
Ding, Zifeng
Yan, Sikuan
Yuan, Zhangdie
Hu, Xianglong
Lin, Fangru
Vlachos, Andreas
Artificial Intelligence
Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity. To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark that jointly assesses both capabilities. Each instance in TCP features a naturalistic dialogue around a collaborative project, where diverse and interdependent temporal constraints are explicitly or implicitly expressed, and models must infer an optimal schedule that satisfies all constraints. To construct TCP, we generate abstract problem prototypes that are then paired with realistic scenarios from various domains and enriched into dialogues using an LLM. A human quality check is performed on a sampled subset to confirm the reliability of our benchmark. We evaluate state-of-the-art LLMs and find that even the strongest models may struggle with TCP, highlighting its difficulty and revealing limitations in LLMs' temporal constraint-based planning abilities. We analyze underlying failure cases, open source our benchmark, and hope our findings can inspire future research.
title TCP: a Benchmark for Temporal Constraint-Based Planning
topic Artificial Intelligence
url https://arxiv.org/abs/2505.19927