A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
Fuente:
arXiv
Saved in:
| Main Authors: | Madge, Chris, Poesio, Massimo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models as Minecraft Agents
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
by: Gan, Yujian, et al.
Published: (2024)
by: Gan, Yujian, et al.
Published: (2024)
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
by: Shao, Juexi, et al.
Published: (2025)
by: Shao, Juexi, et al.
Published: (2025)
Nebula: A discourse aware Minecraft Builder
by: Chaturvedi, Akshay, et al.
Published: (2024)
by: Chaturvedi, Akshay, et al.
Published: (2024)
Integrating knowledge bases to improve coreference and bridging resolution for the chemical domain
by: Lu, Pengcheng, et al.
Published: (2024)
by: Lu, Pengcheng, et al.
Published: (2024)
Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTask
by: Li, Nan, et al.
Published: (2025)
by: Li, Nan, et al.
Published: (2025)
Data Augmentation for Fake Reviews Detection in Multiple Languages and Multiple Domains
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Understanding The Effect Of Temperature On Alignment With Human Opinions
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
BAR: A Backward Reasoning based Agent for Complex Minecraft Tasks
by: Du, Weihong, et al.
Published: (2025)
by: Du, Weihong, et al.
Published: (2025)
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
by: Pavlovic, Maja, et al.
Published: (2024)
by: Pavlovic, Maja, et al.
Published: (2024)
Can LLMs Detect Ambiguous Plural Reference? An Analysis of Split-Antecedent and Mereological Reference
by: Anh, Dang, et al.
Published: (2025)
by: Anh, Dang, et al.
Published: (2025)
An Assessment of Human vs. Model Uncertainty in Soft-Label Learning and Calibration
by: Pavlovic, Maja, et al.
Published: (2026)
by: Pavlovic, Maja, et al.
Published: (2026)
Collaborative Quest Completion with LLM-driven Non-Player Characters in Minecraft
by: Rao, Sudha, et al.
Published: (2024)
by: Rao, Sudha, et al.
Published: (2024)
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
by: Long, Qian, et al.
Published: (2024)
by: Long, Qian, et al.
Published: (2024)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
by: Baidya, Avinash, et al.
Published: (2025)
by: Baidya, Avinash, et al.
Published: (2025)
Talking-to-Build: How LLM-Assisted Interface Shapes Player Performance and Experience in Minecraft
by: Sun, Xin, et al.
Published: (2025)
by: Sun, Xin, et al.
Published: (2025)
DARD: A Multi-Agent Approach for Task-Oriented Dialog Systems
by: Gupta, Aman, et al.
Published: (2024)
by: Gupta, Aman, et al.
Published: (2024)
Improving LLMs' Learning for Coreference Resolution
by: Gan, Yujian, et al.
Published: (2025)
by: Gan, Yujian, et al.
Published: (2025)
Human Label Variation in Implicit Discourse Relation Recognition
by: Yung, Frances, et al.
Published: (2026)
by: Yung, Frances, et al.
Published: (2026)
Extending Activation Steering to Broad Skills and Multiple Behaviours
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks
by: Xu, Frank F., et al.
Published: (2024)
by: Xu, Frank F., et al.
Published: (2024)
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
MCPDial: A Minecraft Persona-driven Dialogue Dataset
by: Alavi, Seyed Hossein, et al.
Published: (2024)
by: Alavi, Seyed Hossein, et al.
Published: (2024)
Evaluating and Enhancing Out-of-Domain Generalization of Task-Oriented Dialog Systems for Task Completion without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2025)
by: Mosharrof, Adib, et al.
Published: (2025)
Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
BAP v2: An Enhanced Task Framework for Instruction Following in Minecraft Dialogues
by: Jayannavar, Prashant, et al.
Published: (2025)
by: Jayannavar, Prashant, et al.
Published: (2025)
Actionable Conversational Quality Indicators for Improving Task-Oriented Dialog Systems
by: Higgins, Michael, et al.
Published: (2021)
by: Higgins, Michael, et al.
Published: (2021)
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
by: Leonardelli, Elisa, et al.
Published: (2025)
by: Leonardelli, Elisa, et al.
Published: (2025)
Training Zero-Shot Generalizable End-to-End Task-Oriented Dialog System Without Turn-level Dialog Annotations
by: Mosharrof, Adib, et al.
Published: (2024)
by: Mosharrof, Adib, et al.
Published: (2024)
AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
by: Liu, Xuannan, et al.
Published: (2026)
by: Liu, Xuannan, et al.
Published: (2026)
TOAD: Task-Oriented Automatic Dialogs with Diverse Response Styles
by: Liu, Yinhong, et al.
Published: (2024)
by: Liu, Yinhong, et al.
Published: (2024)
Dialog Flow Induction for Constrainable LLM-Based Chatbots
by: Agrawal, Stuti, et al.
Published: (2024)
by: Agrawal, Stuti, et al.
Published: (2024)
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
by: Rabbani, Parisa, et al.
Published: (2026)
by: Rabbani, Parisa, et al.
Published: (2026)
Afrispeech-Dialog: A Benchmark Dataset for Spontaneous English Conversations in Healthcare and Beyond
by: Sanni, Mardhiyah, et al.
Published: (2025)
by: Sanni, Mardhiyah, et al.
Published: (2025)
Conversation Routines: A Prompt Engineering Framework for Task-Oriented Dialog Systems
by: Robino, Giorgio
Published: (2025)
by: Robino, Giorgio
Published: (2025)
VAL: Interactive Task Learning with GPT Dialog Parsing
by: Lawley, Lane, et al.
Published: (2023)
by: Lawley, Lane, et al.
Published: (2023)
Synergizing In-context Learning with Hints for End-to-end Task-oriented Dialog Systems
by: Saley, Vishal Vivek, et al.
Published: (2024)
by: Saley, Vishal Vivek, et al.
Published: (2024)
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
by: Yu, Zijian, et al.
Published: (2026)
by: Yu, Zijian, et al.
Published: (2026)
Similar Items
-
Large Language Models as Minecraft Agents
by: Madge, Chris, et al.
Published: (2024) -
Referential ambiguity and clarification requests: comparing human and LLM behaviour
by: Madge, Chris, et al.
Published: (2025) -
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025) -
ClarQ-LLM: A Benchmark for Models Clarifying and Requesting Information in Task-Oriented Dialog
by: Gan, Yujian, et al.
Published: (2024) -
Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
by: Shao, Juexi, et al.
Published: (2025)