Playpen: An Environment for Exploring Learning Through Conversational Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Horst, Nicola, Mazzaccara, Davide, Schmidt, Antonia, Sullivan, Michael, Momentè, Filippo, Franceschetti, Luca, Sadler, Philipp, Hakimov, Sherzod, Testoni, Alberto, Bernardi, Raffaella, Fernández, Raquel, Koller, Alexander, Lemon, Oliver, Schlangen, David, Giulianelli, Mario, Suglia, Alessandro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
by: Momentè, Filippo, et al.
Published: (2025)
by: Momentè, Filippo, et al.
Published: (2025)
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
by: Mazzaccara, Davide, et al.
Published: (2024)
by: Mazzaccara, Davide, et al.
Published: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
by: Sadler, Philipp, et al.
Published: (2024)
by: Sadler, Philipp, et al.
Published: (2024)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
by: Sadler, Philipp, et al.
Published: (2024)
by: Sadler, Philipp, et al.
Published: (2024)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
by: Jordan, Jonathan, et al.
Published: (2025)
by: Jordan, Jonathan, et al.
Published: (2025)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
by: Schlangen, David, et al.
Published: (2025)
by: Schlangen, David, et al.
Published: (2025)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
by: Kranti, Chalamalasetti, et al.
Published: (2026)
by: Kranti, Chalamalasetti, et al.
Published: (2026)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
by: Kranti, Chalamalasetti, et al.
Published: (2024)
by: Kranti, Chalamalasetti, et al.
Published: (2024)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
by: Bhavsar, Nidhir, et al.
Published: (2024)
by: Bhavsar, Nidhir, et al.
Published: (2024)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
by: Hakimov, Sherzod, et al.
Published: (2024)
by: Hakimov, Sherzod, et al.
Published: (2024)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
by: Beyer, Anne, et al.
Published: (2024)
by: Beyer, Anne, et al.
Published: (2024)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
by: Hakimov, Sherzod, et al.
Published: (2025)
by: Hakimov, Sherzod, et al.
Published: (2025)
TurkicNLP: An NLP Toolkit for Turkic Languages
by: Hakimov, Sherzod
Published: (2026)
by: Hakimov, Sherzod
Published: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
by: Hakimov, Sherzod, et al.
Published: (2026)
by: Hakimov, Sherzod, et al.
Published: (2026)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
by: Hakimov, Sherzod, et al.
Published: (2023)
by: Hakimov, Sherzod, et al.
Published: (2023)
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
by: Bavaresco, Anna, et al.
Published: (2024)
by: Bavaresco, Anna, et al.
Published: (2024)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
by: Thakkar, Gaurish, et al.
Published: (2024)
by: Thakkar, Gaurish, et al.
Published: (2024)
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization
by: Borec, Luka, et al.
Published: (2024)
by: Borec, Luka, et al.
Published: (2024)
Free-text Rationale Generation under Readability Level Control
by: Hsu, Yi-Sheng, et al.
Published: (2024)
by: Hsu, Yi-Sheng, et al.
Published: (2024)
A Dialogue Game for Eliciting Balanced Collaboration
by: Jeknić, Isidora, et al.
Published: (2024)
by: Jeknić, Isidora, et al.
Published: (2024)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Procedural Environment Generation for Tool-Use Agents
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
Apresentação exuberante de caso de esclerose sistêmica
by: Gabriela Momente Miquelin
Published: (2018)
by: Gabriela Momente Miquelin
Published: (2018)
Estudo comparativo do uso da minociclina sistêmica versus corticoterapia sistêmica no tratamento de vitiligo em atividade
by: Gabriela Momente Miquelin
Published: (2019)
by: Gabriela Momente Miquelin
Published: (2019)
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
PERMANÊNCIAS E RUPTURAS: SENTIDOS DE GÊNERO EM MULHERES CHEFES DE FAMÍLIA
by: Raquel Jaqueline Freiberger Testoni
Published: (2006)
by: Raquel Jaqueline Freiberger Testoni
Published: (2006)
La Red ¿democrática? en la sociedad del aislamiento
by: Santiago Giulianelli
Published: (2005)
by: Santiago Giulianelli
Published: (2005)
Accelerating Constrained Decoding with Token Space Compression
by: Sullivan, Michael, et al.
Published: (2026)
by: Sullivan, Michael, et al.
Published: (2026)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
Explaining Russian-German code-mixing
by: Hakimov, Nikolay
Published: (2022)
by: Hakimov, Nikolay
Published: (2022)
BO'YIN GRIJALARI BILAN KASALLANGAN BEMORLARNI REABILITATSIYA QILISH VA FIZIOTERAPIYA PROTOKOLLARI
by: Hakimov, Zohidjon
Published: (2025)
by: Hakimov, Zohidjon
Published: (2025)
SHAXSNING O'Z-O'ZIGA ISHONCHI VA UNING TEMPERAMENTGA BOG'LIQ XUSUSIYATLARI.
by: Hakimov, Jamshid
Published: (2026)
by: Hakimov, Jamshid
Published: (2026)
RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
by: Testoni, Alberto, et al.
Published: (2024)
by: Testoni, Alberto, et al.
Published: (2024)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
by: Yang, Jing, et al.
Published: (2026)
by: Yang, Jing, et al.
Published: (2026)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
by: Schlangen, David
Published: (2024)
by: Schlangen, David
Published: (2024)
Similar Items
-
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests
by: Momentè, Filippo, et al.
Published: (2025) -
Learning to Ask Informative Questions: Enhancing LLMs with Preference Optimization and Expected Information Gain
by: Mazzaccara, Davide, et al.
Published: (2024) -
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
by: Sadler, Philipp, et al.
Published: (2024) -
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
by: Sadler, Philipp, et al.
Published: (2024) -
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
by: Kranti, Chalamalasetti, et al.
Published: (2024)