Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
Fuente:
arXiv
Salvato in:
| Autori principali: | Jordan, Jonathan, Hakimov, Sherzod, Schlangen, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024)
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
di: Schlangen, David, et al.
Pubblicazione: (2025)
di: Schlangen, David, et al.
Pubblicazione: (2025)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2026)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2026)
TurkicNLP: An NLP Toolkit for Turkic Languages
di: Hakimov, Sherzod
Pubblicazione: (2026)
di: Hakimov, Sherzod
Pubblicazione: (2026)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
di: Beyer, Anne, et al.
Pubblicazione: (2024)
di: Beyer, Anne, et al.
Pubblicazione: (2024)
Incremental Processing in the Age of Non-Incremental Encoders: An Empirical Assessment of Bidirectional Models for Incremental NLU
di: Madureira, Brielen, et al.
Pubblicazione: (2020)
di: Madureira, Brielen, et al.
Pubblicazione: (2020)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2024)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2024)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
di: Hakimov, Sherzod, et al.
Pubblicazione: (2023)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2023)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
di: Thakkar, Gaurish, et al.
Pubblicazione: (2024)
di: Thakkar, Gaurish, et al.
Pubblicazione: (2024)
Towards Incremental Transformers: An Empirical Analysis of Transformer Models for Incremental NLU
di: Kahardipraja, Patrick, et al.
Pubblicazione: (2021)
di: Kahardipraja, Patrick, et al.
Pubblicazione: (2021)
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
di: Kennington, Casey, et al.
Pubblicazione: (2025)
di: Kennington, Casey, et al.
Pubblicazione: (2025)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
di: Madureira, Brielen, et al.
Pubblicazione: (2022)
di: Madureira, Brielen, et al.
Pubblicazione: (2022)
Free-text Rationale Generation under Readability Level Control
di: Hsu, Yi-Sheng, et al.
Pubblicazione: (2024)
di: Hsu, Yi-Sheng, et al.
Pubblicazione: (2024)
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
di: Lu, Christina, et al.
Pubblicazione: (2026)
di: Lu, Christina, et al.
Pubblicazione: (2026)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
di: Horst, Nicola, et al.
Pubblicazione: (2025)
di: Horst, Nicola, et al.
Pubblicazione: (2025)
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation
di: Diallo, Aissatou, et al.
Pubblicazione: (2025)
di: Diallo, Aissatou, et al.
Pubblicazione: (2025)
Large Language Model based Situational Dialogues for Second Language Learning
di: Xu, Shuyao, et al.
Pubblicazione: (2024)
di: Xu, Shuyao, et al.
Pubblicazione: (2024)
Situated Natural Language Explanations
di: Zhu, Zining, et al.
Pubblicazione: (2023)
di: Zhu, Zining, et al.
Pubblicazione: (2023)
When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-Incrementality
di: Madureira, Brielen, et al.
Pubblicazione: (2024)
di: Madureira, Brielen, et al.
Pubblicazione: (2024)
It Couldn't Help But Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised Learning
di: Madureira, Brielen, et al.
Pubblicazione: (2024)
di: Madureira, Brielen, et al.
Pubblicazione: (2024)
Characterizing Language Use in a Collaborative Situated Game
di: Tomlin, Nicholas, et al.
Pubblicazione: (2025)
di: Tomlin, Nicholas, et al.
Pubblicazione: (2025)
"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration
di: Bohus, Dan, et al.
Pubblicazione: (2024)
di: Bohus, Dan, et al.
Pubblicazione: (2024)
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization
di: Borec, Luka, et al.
Pubblicazione: (2024)
di: Borec, Luka, et al.
Pubblicazione: (2024)
A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
di: Decostanzi, Ivan, et al.
Pubblicazione: (2025)
di: Decostanzi, Ivan, et al.
Pubblicazione: (2025)
Beyond Static Personas: Situational Personality Steering for Large Language Models
di: Wei, Zesheng, et al.
Pubblicazione: (2026)
di: Wei, Zesheng, et al.
Pubblicazione: (2026)
Mars: Situated Inductive Reasoning in an Open-World Environment
di: Tang, Xiaojuan, et al.
Pubblicazione: (2024)
di: Tang, Xiaojuan, et al.
Pubblicazione: (2024)
SPRI: Aligning Large Language Models with Context-Situated Principles
di: Zhan, Hongli, et al.
Pubblicazione: (2025)
di: Zhan, Hongli, et al.
Pubblicazione: (2025)
Solving Situation Puzzles with Large Language Model and External Reformulation
di: Li, Kun, et al.
Pubblicazione: (2025)
di: Li, Kun, et al.
Pubblicazione: (2025)
Scenarios and Approaches for Situated Natural Language Explanations
di: Qiu, Pengshuo, et al.
Pubblicazione: (2024)
di: Qiu, Pengshuo, et al.
Pubblicazione: (2024)
Multimodal Situational Safety
di: Zhou, Kaiwen, et al.
Pubblicazione: (2024)
di: Zhou, Kaiwen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024) -
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024) -
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025) -
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025) -
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
di: Sadler, Philipp, et al.
Pubblicazione: (2024)