clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
Fuente:
arXiv
Guardado en:
| Autores principales: | Kranti, Chalamalasetti, Hakimov, Sherzod, Schlangen, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
por: Kranti, Chalamalasetti, et al.
Publicado: (2026)
por: Kranti, Chalamalasetti, et al.
Publicado: (2026)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
por: Schlangen, David, et al.
Publicado: (2025)
por: Schlangen, David, et al.
Publicado: (2025)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
por: Beyer, Anne, et al.
Publicado: (2024)
por: Beyer, Anne, et al.
Publicado: (2024)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
por: Jordan, Jonathan, et al.
Publicado: (2025)
por: Jordan, Jonathan, et al.
Publicado: (2025)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
por: Hakimov, Sherzod, et al.
Publicado: (2025)
por: Hakimov, Sherzod, et al.
Publicado: (2025)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
por: Sadler, Philipp, et al.
Publicado: (2024)
por: Sadler, Philipp, et al.
Publicado: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
por: Sadler, Philipp, et al.
Publicado: (2024)
por: Sadler, Philipp, et al.
Publicado: (2024)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
por: Hakimov, Sherzod, et al.
Publicado: (2026)
por: Hakimov, Sherzod, et al.
Publicado: (2026)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
por: Bhavsar, Nidhir, et al.
Publicado: (2024)
por: Bhavsar, Nidhir, et al.
Publicado: (2024)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
TurkicNLP: An NLP Toolkit for Turkic Languages
por: Hakimov, Sherzod
Publicado: (2026)
por: Hakimov, Sherzod
Publicado: (2026)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
por: Kranti, Chalamalasetti
Publicado: (2025)
por: Kranti, Chalamalasetti
Publicado: (2025)
Test Set Quality in Multilingual LLM Evaluation
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
por: Hakimov, Sherzod, et al.
Publicado: (2023)
por: Hakimov, Sherzod, et al.
Publicado: (2023)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
por: Hakimov, Sherzod, et al.
Publicado: (2024)
por: Hakimov, Sherzod, et al.
Publicado: (2024)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
por: Thakkar, Gaurish, et al.
Publicado: (2024)
por: Thakkar, Gaurish, et al.
Publicado: (2024)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
por: Madureira, Brielen, et al.
Publicado: (2022)
por: Madureira, Brielen, et al.
Publicado: (2022)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
por: Hakimov, Sherzod, et al.
Publicado: (2025)
por: Hakimov, Sherzod, et al.
Publicado: (2025)
Free-text Rationale Generation under Readability Level Control
por: Hsu, Yi-Sheng, et al.
Publicado: (2024)
por: Hsu, Yi-Sheng, et al.
Publicado: (2024)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
por: Zhang, Yifei, et al.
Publicado: (2026)
por: Zhang, Yifei, et al.
Publicado: (2026)
A Dialogue Game for Eliciting Balanced Collaboration
por: Jeknić, Isidora, et al.
Publicado: (2024)
por: Jeknić, Isidora, et al.
Publicado: (2024)
Spec-TOD: A Specialized Instruction-Tuned LLM Framework for Efficient Task-Oriented Dialogue Systems
por: Nguyen, Quang-Vinh, et al.
Publicado: (2025)
por: Nguyen, Quang-Vinh, et al.
Publicado: (2025)
Reliable LLM-based User Simulator for Task-Oriented Dialogue Systems
por: Sekulić, Ivan, et al.
Publicado: (2024)
por: Sekulić, Ivan, et al.
Publicado: (2024)
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
por: Kennington, Casey, et al.
Publicado: (2025)
por: Kennington, Casey, et al.
Publicado: (2025)
CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment
por: Shayanfar, Radin, et al.
Publicado: (2025)
por: Shayanfar, Radin, et al.
Publicado: (2025)
TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
por: Ghazarian, Sarik, et al.
Publicado: (2025)
por: Ghazarian, Sarik, et al.
Publicado: (2025)
Task-Oriented Dialogue with In-Context Learning
por: Bocklisch, Tom, et al.
Publicado: (2024)
por: Bocklisch, Tom, et al.
Publicado: (2024)
Bridging Reasoning and Action: Hybrid LLM-RL Framework for Efficient Cross-Domain Task-Oriented Dialogue
por: Zhao, Yangyang, et al.
Publicado: (2026)
por: Zhao, Yangyang, et al.
Publicado: (2026)
DiagGPT: An LLM-based and Multi-agent Dialogue System with Automatic Topic Management for Flexible Task-Oriented Dialogue
por: Cao, Lang
Publicado: (2023)
por: Cao, Lang
Publicado: (2023)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
por: Yang, Jing, et al.
Publicado: (2026)
por: Yang, Jing, et al.
Publicado: (2026)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
por: Schlangen, David
Publicado: (2024)
por: Schlangen, David
Publicado: (2024)
CAUSE: Counterfactual Assessment of User Satisfaction Estimation in Task-Oriented Dialogue Systems
por: Abolghasemi, Amin, et al.
Publicado: (2024)
por: Abolghasemi, Amin, et al.
Publicado: (2024)
Bootstrapping LLM-based Task-Oriented Dialogue Agents via Self-Talk
por: Ulmer, Dennis, et al.
Publicado: (2024)
por: Ulmer, Dennis, et al.
Publicado: (2024)
LLM-Driven Multi-Turn Task-Oriented Dialogue Synthesis for Realistic Reasoning
por: Zhu, Yu, et al.
Publicado: (2026)
por: Zhu, Yu, et al.
Publicado: (2026)
HierTOD: A Task-Oriented Dialogue System Driven by Hierarchical Goals
por: Mo, Lingbo, et al.
Publicado: (2024)
por: Mo, Lingbo, et al.
Publicado: (2024)
Simulating User Diversity in Task-Oriented Dialogue Systems using Large Language Models
por: Ahmad, Adnan, et al.
Publicado: (2025)
por: Ahmad, Adnan, et al.
Publicado: (2025)
Ejemplares similares
-
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
por: Kranti, Chalamalasetti, et al.
Publicado: (2026) -
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
por: Kranti, Chalamalasetti, et al.
Publicado: (2024) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
por: Kranti, Chalamalasetti, et al.
Publicado: (2025) -
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
por: Kranti, Chalamalasetti, et al.
Publicado: (2024) -
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
por: Schlangen, David, et al.
Publicado: (2025)