A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schlangen, David, Hakimov, Sherzod, Kranti, Chalamalasetti, Jordan, Jonathan, Sadler, Philipp |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
von: Sadler, Philipp, et al.
Veröffentlicht: (2024)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
von: Jordan, Jonathan, et al.
Veröffentlicht: (2025)
von: Jordan, Jonathan, et al.
Veröffentlicht: (2025)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
von: Bhavsar, Nidhir, et al.
Veröffentlicht: (2024)
von: Bhavsar, Nidhir, et al.
Veröffentlicht: (2024)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2026)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
von: Kranti, Chalamalasetti
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti
Veröffentlicht: (2025)
Test Set Quality in Multilingual LLM Evaluation
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
TurkicNLP: An NLP Toolkit for Turkic Languages
von: Hakimov, Sherzod
Veröffentlicht: (2026)
von: Hakimov, Sherzod
Veröffentlicht: (2026)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2024)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2024)
The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization
von: Borec, Luka, et al.
Veröffentlicht: (2024)
von: Borec, Luka, et al.
Veröffentlicht: (2024)
A Dialogue Game for Eliciting Balanced Collaboration
von: Jeknić, Isidora, et al.
Veröffentlicht: (2024)
von: Jeknić, Isidora, et al.
Veröffentlicht: (2024)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2023)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2023)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
von: Thakkar, Gaurish, et al.
Veröffentlicht: (2024)
von: Thakkar, Gaurish, et al.
Veröffentlicht: (2024)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
von: Madureira, Brielen, et al.
Veröffentlicht: (2022)
von: Madureira, Brielen, et al.
Veröffentlicht: (2022)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
von: Schlangen, David
Veröffentlicht: (2024)
von: Schlangen, David
Veröffentlicht: (2024)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025)
von: Hakimov, Sherzod, et al.
Veröffentlicht: (2025)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
von: Yang, Jing, et al.
Veröffentlicht: (2026)
von: Yang, Jing, et al.
Veröffentlicht: (2026)
Free-text Rationale Generation under Readability Level Control
von: Hsu, Yi-Sheng, et al.
Veröffentlicht: (2024)
von: Hsu, Yi-Sheng, et al.
Veröffentlicht: (2024)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
von: Horst, Nicola, et al.
Veröffentlicht: (2025)
von: Horst, Nicola, et al.
Veröffentlicht: (2025)
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
von: Kennington, Casey, et al.
Veröffentlicht: (2025)
von: Kennington, Casey, et al.
Veröffentlicht: (2025)
LinguaGame: A Linguistically Grounded Game-Theoretic Paradigm for Multi-Agent Dialogue Generation
von: Ye, Yuxiao, et al.
Veröffentlicht: (2026)
von: Ye, Yuxiao, et al.
Veröffentlicht: (2026)
How Stylistic Similarity Shapes Preferences in Dialogue Dataset with User and Third Party Evaluations
von: Numaya, Ikumi, et al.
Veröffentlicht: (2025)
von: Numaya, Ikumi, et al.
Veröffentlicht: (2025)
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
von: Momentè, Filippo, et al.
Veröffentlicht: (2025)
Could the Road to Grounded, Neuro-symbolic AI be Paved with Words-as-Classifiers?
von: Kennington, Casey, et al.
Veröffentlicht: (2025)
von: Kennington, Casey, et al.
Veröffentlicht: (2025)
Incremental Processing in the Age of Non-Incremental Encoders: An Empirical Assessment of Bidirectional Models for Incremental NLU
von: Madureira, Brielen, et al.
Veröffentlicht: (2020)
von: Madureira, Brielen, et al.
Veröffentlicht: (2020)
Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
It Couldn't Help But Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised Learning
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
LLM as a Scorer: The Impact of Output Order on Dialogue Evaluation
von: Chen, Yi-Pei, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Pei, et al.
Veröffentlicht: (2024)
When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-Incrementality
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
von: Madureira, Brielen, et al.
Veröffentlicht: (2024)
Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU
von: Kahardipraja, Patrick, et al.
Veröffentlicht: (2021)
von: Kahardipraja, Patrick, et al.
Veröffentlicht: (2021)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
von: Tang, Yuqi, et al.
Veröffentlicht: (2025)
von: Tang, Yuqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025) -
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
von: Beyer, Anne, et al.
Veröffentlicht: (2024) -
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026) -
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)