LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
Fuente:
arXiv
Guardado en:
| Autores principales: | Mantenoglou, Periklis, Hazra, Rishi, Martires, Pedro Zuidberg Dos, De Raedt, Luc |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SayCanPay: Heuristic Planning with Large Language Models using Learnable Domain Knowledge
por: Hazra, Rishi, et al.
Publicado: (2023)
por: Hazra, Rishi, et al.
Publicado: (2023)
Two Constraint Compilation Methods for Lifted Planning
por: Mantenoglou, Periklis, et al.
Publicado: (2025)
por: Mantenoglou, Periklis, et al.
Publicado: (2025)
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
por: Hazra, Rishi, et al.
Publicado: (2025)
por: Hazra, Rishi, et al.
Publicado: (2025)
Can Large Language Models Reason? A Characterization via 3-SAT
por: Hazra, Rishi, et al.
Publicado: (2024)
por: Hazra, Rishi, et al.
Publicado: (2024)
Declarative Probabilistic Logic Programming in Discrete-Continuous Domains
por: Martires, Pedro Zuidberg Dos, et al.
Publicado: (2023)
por: Martires, Pedro Zuidberg Dos, et al.
Publicado: (2023)
Neurosymbolic Decision Trees
por: Möller, Matthias, et al.
Publicado: (2025)
por: Möller, Matthias, et al.
Publicado: (2025)
Probabilistic Neural Circuits
por: Martires, Pedro Zuidberg Dos
Publicado: (2024)
por: Martires, Pedro Zuidberg Dos
Publicado: (2024)
REvolve: Reward Evolution with Large Language Models using Human Feedback
por: Hazra, Rishi, et al.
Publicado: (2024)
por: Hazra, Rishi, et al.
Publicado: (2024)
Semirings for Probabilistic and Neuro-Symbolic Logic Programming
por: Derkinderen, Vincent, et al.
Publicado: (2024)
por: Derkinderen, Vincent, et al.
Publicado: (2024)
Efficient Temporal Datalog Materialisation for Composite Event Recognition
por: Mantenoglou, Periklis
Publicado: (2026)
por: Mantenoglou, Periklis
Publicado: (2026)
A Quantum Information Theoretic Approach to Tractable Probabilistic Models
por: Martires, Pedro Zuidberg Dos
Publicado: (2025)
por: Martires, Pedro Zuidberg Dos
Publicado: (2025)
COvolve: Adversarial Co-Evolution of Large-Language-Model-Generated Policies and Environments via Two-Player Zero-Sum Game
por: Sygkounas, Alkis, et al.
Publicado: (2026)
por: Sygkounas, Alkis, et al.
Publicado: (2026)
Independence Is Not an Issue in Neurosymbolic AI
por: Faronius, Håkan Karlsson, et al.
Publicado: (2025)
por: Faronius, Håkan Karlsson, et al.
Publicado: (2025)
A Fast Convoluted Story: Scaling Probabilistic Inference for Integer Arithmetic
por: De Smet, Lennert, et al.
Publicado: (2024)
por: De Smet, Lennert, et al.
Publicado: (2024)
Automated Reasoning in Systems Biology: a Necessity for Precision Medicine
por: Martires, Pedro Zuidberg Dos, et al.
Publicado: (2024)
por: Martires, Pedro Zuidberg Dos, et al.
Publicado: (2024)
Valid Text-to-SQL Generation with Unification-based DeepStochLog
por: Jiao, Ying, et al.
Publicado: (2025)
por: Jiao, Ying, et al.
Publicado: (2025)
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
por: Zheng, Huaixiu Steven, et al.
Publicado: (2024)
por: Zheng, Huaixiu Steven, et al.
Publicado: (2024)
Building Safe and Deployable Clinical Natural Language Processing under Temporal Leakage Constraints
por: Cho, Ha Na, et al.
Publicado: (2026)
por: Cho, Ha Na, et al.
Publicado: (2026)
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
por: Zhang, Yinger, et al.
Publicado: (2026)
por: Zhang, Yinger, et al.
Publicado: (2026)
ChinaTravel: An Open-Ended Travel Planning Benchmark with Compositional Constraint Validation for Language Agents
por: Shao, Jie-Jing, et al.
Publicado: (2024)
por: Shao, Jie-Jing, et al.
Publicado: (2024)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
por: Chen, Pei-An, et al.
Publicado: (2026)
por: Chen, Pei-An, et al.
Publicado: (2026)
APC-RL: Exceeding Data-Driven Behavior Priors with Adaptive Policy Composition
por: Rietz, Finn, et al.
Publicado: (2026)
por: Rietz, Finn, et al.
Publicado: (2026)
To Tell The Truth: Language of Deception and Language Models
por: Hazra, Sanchaita, et al.
Publicado: (2023)
por: Hazra, Sanchaita, et al.
Publicado: (2023)
CCTU: A Benchmark for Tool Use under Complex Constraints
por: Ye, Junjie, et al.
Publicado: (2026)
por: Ye, Junjie, et al.
Publicado: (2026)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
por: Kammakomati, Mehant, et al.
Publicado: (2024)
por: Kammakomati, Mehant, et al.
Publicado: (2024)
TripCraft: A Benchmark for Spatio-Temporally Fine Grained Travel Planning
por: Chaudhuri, Soumyabrata, et al.
Publicado: (2025)
por: Chaudhuri, Soumyabrata, et al.
Publicado: (2025)
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
por: Banerjee, Somnath, et al.
Publicado: (2025)
por: Banerjee, Somnath, et al.
Publicado: (2025)
TripTide: A Benchmark for Adaptive Travel Planning under Disruptions
por: Karmakar, Priyanshu, et al.
Publicado: (2025)
por: Karmakar, Priyanshu, et al.
Publicado: (2025)
Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs
por: English, William, et al.
Publicado: (2025)
por: English, William, et al.
Publicado: (2025)
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
por: Roccabruna, Gabriel, et al.
Publicado: (2026)
por: Roccabruna, Gabriel, et al.
Publicado: (2026)
UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models
por: Zheng, Yu, et al.
Publicado: (2025)
por: Zheng, Yu, et al.
Publicado: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
por: Tang, Weizhi, et al.
Publicado: (2024)
por: Tang, Weizhi, et al.
Publicado: (2024)
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
por: Bhatia, Gagan, et al.
Publicado: (2024)
por: Bhatia, Gagan, et al.
Publicado: (2024)
Large Language Model Meets Constraint Propagation
por: Bonlarron, Alexandre, et al.
Publicado: (2025)
por: Bonlarron, Alexandre, et al.
Publicado: (2025)
GinSign: Grounding Natural Language Into System Signatures for Temporal Logic Translation
por: English, William, et al.
Publicado: (2025)
por: English, William, et al.
Publicado: (2025)
GraphicBench: A Planning Benchmark for Graphic Design with Language Agents
por: Ki, Dayeon, et al.
Publicado: (2025)
por: Ki, Dayeon, et al.
Publicado: (2025)
Empirical Characterization of Temporal Constraint Processing in LLMs
por: Marín, Javier
Publicado: (2025)
por: Marín, Javier
Publicado: (2025)
Planning with Multi-Constraints via Collaborative Language Agents
por: Zhang, Cong, et al.
Publicado: (2024)
por: Zhang, Cong, et al.
Publicado: (2024)
Combining Constraint Programming Reasoning with Large Language Model Predictions
por: Régin, Florian, et al.
Publicado: (2024)
por: Régin, Florian, et al.
Publicado: (2024)
ConDABench: Interactive Evaluation of Language Models for Data Analysis
por: Dutta, Avik, et al.
Publicado: (2025)
por: Dutta, Avik, et al.
Publicado: (2025)
Ejemplares similares
-
SayCanPay: Heuristic Planning with Large Language Models using Learnable Domain Knowledge
por: Hazra, Rishi, et al.
Publicado: (2023) -
Two Constraint Compilation Methods for Lifted Planning
por: Mantenoglou, Periklis, et al.
Publicado: (2025) -
Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
por: Hazra, Rishi, et al.
Publicado: (2025) -
Can Large Language Models Reason? A Characterization via 3-SAT
por: Hazra, Rishi, et al.
Publicado: (2024) -
Declarative Probabilistic Logic Programming in Discrete-Continuous Domains
por: Martires, Pedro Zuidberg Dos, et al.
Publicado: (2023)