clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Beyer, Anne, Chalamalasetti, Kranti, Hakimov, Sherzod, Madureira, Brielen, Sadler, Philipp, Schlangen, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025)
by: Schlangen, David, et al.
Published: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026)
by: Kranti, Chalamalasetti, et al.
Published: (2026)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
by: Sadler, Philipp, et al.
Published: (2024)
by: Sadler, Philipp, et al.
Published: (2024)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
by: Sadler, Philipp, et al.
Published: (2024)
by: Sadler, Philipp, et al.
Published: (2024)
Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests
by: Madureira, Brielen, et al.
Published: (2024)
by: Madureira, Brielen, et al.
Published: (2024)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
by: Madureira, Brielen, et al.
Published: (2022)
by: Madureira, Brielen, et al.
Published: (2022)
Incremental Processing in the Age of Non-Incremental Encoders: An Empirical Assessment of Bidirectional Models for Incremental NLU
by: Madureira, Brielen, et al.
Published: (2020)
by: Madureira, Brielen, et al.
Published: (2020)
It Couldn't Help But Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised Learning
by: Madureira, Brielen, et al.
Published: (2024)
by: Madureira, Brielen, et al.
Published: (2024)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
by: Jordan, Jonathan, et al.
Published: (2025)
by: Jordan, Jonathan, et al.
Published: (2025)
When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-Incrementality
by: Madureira, Brielen, et al.
Published: (2024)
by: Madureira, Brielen, et al.
Published: (2024)
Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU
by: Kahardipraja, Patrick, et al.
Published: (2021)
by: Kahardipraja, Patrick, et al.
Published: (2021)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
by: Kranti, Chalamalasetti
Published: (2025)
by: Kranti, Chalamalasetti
Published: (2025)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
by: Hakimov, Sherzod, et al.
Published: (2024)
by: Hakimov, Sherzod, et al.
Published: (2024)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
by: Bhavsar, Nidhir, et al.
Published: (2024)
by: Bhavsar, Nidhir, et al.
Published: (2024)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
by: Hakimov, Sherzod, et al.
Published: (2023)
by: Hakimov, Sherzod, et al.
Published: (2023)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
TurkicNLP: An NLP Toolkit for Turkic Languages
by: Hakimov, Sherzod
Published: (2026)
by: Hakimov, Sherzod
Published: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
by: Hakimov, Sherzod, et al.
Published: (2026)
by: Hakimov, Sherzod, et al.
Published: (2026)
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
by: Thakkar, Gaurish, et al.
Published: (2024)
by: Thakkar, Gaurish, et al.
Published: (2024)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization
by: Borec, Luka, et al.
Published: (2024)
by: Borec, Luka, et al.
Published: (2024)
Geolocating News about Extreme Climate Events: A Comparative Analysis of Off-the-Shelf Tools for Toponym Identification in German
by: Madureira, Brielen, et al.
Published: (2026)
by: Madureira, Brielen, et al.
Published: (2026)
Retrieving Floods without Floodlights: Topic Models as Binary Classifiers for Extreme Climate Events in German News
by: Madureira, Brielen, et al.
Published: (2026)
by: Madureira, Brielen, et al.
Published: (2026)
Free-text Rationale Generation under Readability Level Control
by: Hsu, Yi-Sheng, et al.
Published: (2024)
by: Hsu, Yi-Sheng, et al.
Published: (2024)
The Newsworthiness of Brazilian Distress: A Peak Analysis on Time Series of International Media Attention to Disasters in Brazil
by: Madureira, Brielen, et al.
Published: (2026)
by: Madureira, Brielen, et al.
Published: (2026)
How Loud Rumbles Hit Newsstands: A Data Analysis of Coverage and Spatial Bias in German News about Landslides Around the World
by: Madureira, Brielen, et al.
Published: (2026)
by: Madureira, Brielen, et al.
Published: (2026)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
by: Schlangen, David
Published: (2024)
by: Schlangen, David
Published: (2024)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
by: Horst, Nicola, et al.
Published: (2025)
by: Horst, Nicola, et al.
Published: (2025)
MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
by: Fabbri, Alexander R., et al.
Published: (2025)
by: Fabbri, Alexander R., et al.
Published: (2025)
Explaining Russian-German code-mixing
by: Hakimov, Nikolay
Published: (2022)
by: Hakimov, Nikolay
Published: (2022)
BO'YIN GRIJALARI BILAN KASALLANGAN BEMORLARNI REABILITATSIYA QILISH VA FIZIOTERAPIYA PROTOKOLLARI
by: Hakimov, Zohidjon
Published: (2025)
by: Hakimov, Zohidjon
Published: (2025)
SHAXSNING O'Z-O'ZIGA ISHONCHI VA UNING TEMPERAMENTGA BOG'LIQ XUSUSIYATLARI.
by: Hakimov, Jamshid
Published: (2026)
by: Hakimov, Jamshid
Published: (2026)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
by: He, Yun, et al.
Published: (2024)
by: He, Yun, et al.
Published: (2024)
Similar Items
-
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026) -
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
by: Kranti, Chalamalasetti, et al.
Published: (2025) -
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
by: Kranti, Chalamalasetti, et al.
Published: (2024)