How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
Fuente:
arXiv
Saved in:
| Main Authors: | Bhavsar, Nidhir, Jordan, Jonathan, Hakimov, Sherzod, Schlangen, David |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
by: Jordan, Jonathan, et al.
Published: (2025)
by: Jordan, Jonathan, et al.
Published: (2025)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025)
by: Schlangen, David, et al.
Published: (2025)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
by: Sadler, Philipp, et al.
Published: (2024)
by: Sadler, Philipp, et al.
Published: (2024)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
by: Sadler, Philipp, et al.
Published: (2024)
by: Sadler, Philipp, et al.
Published: (2024)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026)
by: Kranti, Chalamalasetti, et al.
Published: (2026)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
by: Hakimov, Sherzod, et al.
Published: (2024)
by: Hakimov, Sherzod, et al.
Published: (2024)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
by: Hakimov, Sherzod, et al.
Published: (2026)
by: Hakimov, Sherzod, et al.
Published: (2026)
TurkicNLP: An NLP Toolkit for Turkic Languages
by: Hakimov, Sherzod
Published: (2026)
by: Hakimov, Sherzod
Published: (2026)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
by: Beyer, Anne, et al.
Published: (2024)
by: Beyer, Anne, et al.
Published: (2024)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
by: Hakimov, Sherzod, et al.
Published: (2023)
by: Hakimov, Sherzod, et al.
Published: (2023)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
by: Schlangen, David
Published: (2024)
by: Schlangen, David
Published: (2024)
Taking Action Towards Graceful Interaction: The Effects of Performing Actions on Modelling Policies for Instruction Clarification Requests
by: Madureira, Brielen, et al.
Published: (2024)
by: Madureira, Brielen, et al.
Published: (2024)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
by: Thakkar, Gaurish, et al.
Published: (2024)
by: Thakkar, Gaurish, et al.
Published: (2024)
Free-text Rationale Generation under Readability Level Control
by: Hsu, Yi-Sheng, et al.
Published: (2024)
by: Hsu, Yi-Sheng, et al.
Published: (2024)
How Many Bursts Does it Take to Form a Core at the Center of a Galaxy?
by: Mostow, Olivia, et al.
Published: (2024)
by: Mostow, Olivia, et al.
Published: (2024)
Humor in the Academic Library: You Must Be Joking! or, How Many Academic Librarians Does It Take To Change a Lightbulb?
by: Black, Leah, et al.
Published: (1999)
by: Black, Leah, et al.
Published: (1999)
Wiring Switches to Light Bulbs
by: Buckley, Stephen M., et al.
Published: (2011)
by: Buckley, Stephen M., et al.
Published: (2011)
Generative AI at the Crossroads: Light Bulb, Dynamo, or Microscope?
by: Baily, Martin, et al.
Published: (2025)
by: Baily, Martin, et al.
Published: (2025)
How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks
by: Arimbur, Johin Johny
Published: (2026)
by: Arimbur, Johin Johny
Published: (2026)
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
by: Verma, Sahil, et al.
Published: (2024)
by: Verma, Sahil, et al.
Published: (2024)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
by: Horst, Nicola, et al.
Published: (2025)
by: Horst, Nicola, et al.
Published: (2025)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
by: Yang, Jing, et al.
Published: (2026)
by: Yang, Jing, et al.
Published: (2026)
A Dialogue Game for Eliciting Balanced Collaboration
by: Jeknić, Isidora, et al.
Published: (2024)
by: Jeknić, Isidora, et al.
Published: (2024)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
by: Madureira, Brielen, et al.
Published: (2022)
by: Madureira, Brielen, et al.
Published: (2022)
Laser Lights or Dim Bulbs? Evaluating Reference Librarians' Use of Electronic Sources.
by: Welch, Jeanie M.
Published: (1999)
by: Welch, Jeanie M.
Published: (1999)
Aerodynamic Design and Performance Evaluation of Pipe Diffuser for Centrifugal Compressor of Micro Gas Turbine
by: Bhavsar, Sujal, et al.
Published: (2024)
by: Bhavsar, Sujal, et al.
Published: (2024)
How Many Dimensions Does It Take To Measure Users' Perceptions of Libraries?: A LibQUAL+Study.
by: Thompson, Bruce, et al.
Published: (2001)
by: Thompson, Bruce, et al.
Published: (2001)
How Does CEO Cognitive Style and Board Characteristics Impact Strategic Risk‐Taking?
by: Xiaoyu Tian, et al.
Published: (2025)
by: Xiaoyu Tian, et al.
Published: (2025)
Who Plays First? Optimizing the Order of Play in Stackelberg Games with Many Robots
by: Hu, Haimin, et al.
Published: (2024)
by: Hu, Haimin, et al.
Published: (2024)
Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case
by: Chen, Peng, et al.
Published: (2024)
by: Chen, Peng, et al.
Published: (2024)
Game Mining: How to Make Money from those about to Play a Game
by: Bono, James W., et al.
Published: (2024)
by: Bono, James W., et al.
Published: (2024)
Decoding Hypoxia‐Induced Metabolomic Changes in Breast Cancer EVs and Their Functional Effects on Cancer Cells
by: Ashish Sahu, et al.
Published: (2025)
by: Ashish Sahu, et al.
Published: (2025)
Spring-flowering Bulbs and How to Force Them for the Home
by: Howe, Marshall A. (Marshall Avery)
Published: (1924)
by: Howe, Marshall A. (Marshall Avery)
Published: (1924)
Security Concerns in IoT Light Bulbs: Investigating Covert Channels
by: Rohilla, Ravisha, et al.
Published: (2024)
by: Rohilla, Ravisha, et al.
Published: (2024)
Similar Items
-
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
by: Jordan, Jonathan, et al.
Published: (2025) -
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025) -
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025) -
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
by: Sadler, Philipp, et al.
Published: (2024) -
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
by: Sadler, Philipp, et al.
Published: (2024)