Plancraft: an evaluation dataset for planning with LLM agents
Fuente:
arXiv
Saved in:
| Main Authors: | Dagan, Gautier, Keller, Frank, Lascarides, Alex |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$How^{2}$: How to learn from procedural How-to questions
by: Dagan, Gautier, et al.
Published: (2025)
by: Dagan, Gautier, et al.
Published: (2025)
SECURE: Semantics-aware Embodied Conversation under Unawareness for Lifelong Robot Learning
by: Rubavicius, Rimvydas, et al.
Published: (2024)
by: Rubavicius, Rimvydas, et al.
Published: (2024)
Contrastive Learning with Narrative Twins for Modeling Story Salience
by: Sterner, Igor, et al.
Published: (2026)
by: Sterner, Igor, et al.
Published: (2026)
Understanding the planning of LLM agents: A survey
by: Huang, Xu, et al.
Published: (2024)
by: Huang, Xu, et al.
Published: (2024)
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
by: Mosquera, Manuel, et al.
Published: (2024)
by: Mosquera, Manuel, et al.
Published: (2024)
MovieSum: An Abstractive Summarization Dataset for Movie Screenplays
by: Saxena, Rohit, et al.
Published: (2024)
by: Saxena, Rohit, et al.
Published: (2024)
Generating Visual Stories with Grounded and Coreferent Characters
by: Liu, Danyang, et al.
Published: (2024)
by: Liu, Danyang, et al.
Published: (2024)
End-to-End Long Document Summarization using Gradient Caching
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
by: Qu, Yincen, et al.
Published: (2025)
by: Qu, Yincen, et al.
Published: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
by: Panagoulias, Dimitrios P., et al.
Published: (2024)
Diverse And Private Synthetic Datasets Generation for RAG evaluation: A multi-agent framework
by: Driouich, Ilias, et al.
Published: (2025)
by: Driouich, Ilias, et al.
Published: (2025)
Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
by: Swain, Sankalp Tattwadarshi, et al.
Published: (2025)
CrisiText: A dataset of warning messages for LLM training in emergency communication
by: Gonella, Giacomo, et al.
Published: (2025)
by: Gonella, Giacomo, et al.
Published: (2025)
Reliable and diverse evaluation of LLM medical knowledge mastery
by: Zhou, Yuxuan, et al.
Published: (2024)
by: Zhou, Yuxuan, et al.
Published: (2024)
GenerationPrograms: Fine-grained Attribution with Executable Programs
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
MARS: toward more efficient multi-agent collaboration for LLM reasoning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Exploring the features used for summary evaluation by Human and GPT
by: Sadeghi, Zahra, et al.
Published: (2025)
by: Sadeghi, Zahra, et al.
Published: (2025)
AutoHarness: improving LLM agents by automatically synthesizing a code harness
by: Lou, Xinghua, et al.
Published: (2026)
by: Lou, Xinghua, et al.
Published: (2026)
Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
by: Goldman, Omer, et al.
Published: (2024)
by: Goldman, Omer, et al.
Published: (2024)
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
by: Zhang, Shiyue, et al.
Published: (2024)
by: Zhang, Shiyue, et al.
Published: (2024)
Identifying User Goals from UI Trajectories
by: Berkovitch, Omri, et al.
Published: (2024)
by: Berkovitch, Omri, et al.
Published: (2024)
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
by: Harary, Sapir, et al.
Published: (2025)
by: Harary, Sapir, et al.
Published: (2025)
Kalahi: A handcrafted, grassroots cultural LLM evaluation suite for Filipino
by: Montalan, Jann Railey, et al.
Published: (2024)
by: Montalan, Jann Railey, et al.
Published: (2024)
Faithfulness metric fusion: Improving the evaluation of LLM trustworthiness across domains
by: Malin, Ben, et al.
Published: (2025)
by: Malin, Ben, et al.
Published: (2025)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025)
by: Cao, Hongliu, et al.
Published: (2025)
Learning Visually Grounded Domain Ontologies via Embodied Conversation and Explanation
by: Park, Jonghyuk, et al.
Published: (2024)
by: Park, Jonghyuk, et al.
Published: (2024)
Efficient Data Generation for Source-grounded Information-seeking Dialogs: A Use Case for Meeting Transcripts
by: Golany, Lotem, et al.
Published: (2024)
by: Golany, Lotem, et al.
Published: (2024)
Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent
by: Nusrat, Humza, et al.
Published: (2025)
by: Nusrat, Humza, et al.
Published: (2025)
Controlled LLM-based Reasoning for Clinical Trial Retrieval
by: Jullien, Mael, et al.
Published: (2024)
by: Jullien, Mael, et al.
Published: (2024)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
Automated test generation to evaluate tool-augmented LLMs as conversational AI agents
by: Arcadinho, Samuel, et al.
Published: (2024)
by: Arcadinho, Samuel, et al.
Published: (2024)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
by: Ichmoukhamedov, Timour, et al.
Published: (2024)
by: Ichmoukhamedov, Timour, et al.
Published: (2024)
To what extent is ChatGPT useful for language teacher lesson plan creation?
by: Dornburg, Alex, et al.
Published: (2024)
by: Dornburg, Alex, et al.
Published: (2024)
Learn to Disguise: Avoid Refusal Responses in LLM's Defense via a Multi-agent Attacker-Disguiser Game
by: Xu, Qianqiao, et al.
Published: (2024)
by: Xu, Qianqiao, et al.
Published: (2024)
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
by: Jin, Song, et al.
Published: (2025)
by: Jin, Song, et al.
Published: (2025)
Small Models, Big Results: Achieving Superior Intent Extraction through Decomposition
by: Cohen, Danielle, et al.
Published: (2025)
by: Cohen, Danielle, et al.
Published: (2025)
Detecting Errors through Ensembling Prompts (DEEP): An End-to-End LLM Framework for Detecting Factual Errors
by: Chandler, Alex, et al.
Published: (2024)
by: Chandler, Alex, et al.
Published: (2024)
LiTransProQA: an LLM-based Literary Translation evaluation metric with Professional Question Answering
by: Zhang, Ran, et al.
Published: (2025)
by: Zhang, Ran, et al.
Published: (2025)
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
by: Cao, Lang
Published: (2023)
by: Cao, Lang
Published: (2023)
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
by: Tang, Xiuxiu, et al.
Published: (2026)
by: Tang, Xiuxiu, et al.
Published: (2026)
Similar Items
-
$How^{2}$: How to learn from procedural How-to questions
by: Dagan, Gautier, et al.
Published: (2025) -
SECURE: Semantics-aware Embodied Conversation under Unawareness for Lifelong Robot Learning
by: Rubavicius, Rimvydas, et al.
Published: (2024) -
Contrastive Learning with Narrative Twins for Modeling Story Salience
by: Sterner, Igor, et al.
Published: (2026) -
Understanding the planning of LLM agents: A survey
by: Huang, Xu, et al.
Published: (2024) -
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
by: Mosquera, Manuel, et al.
Published: (2024)