Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hakimov, Sherzod, Abdullayeva, Yerkezhan, Koshti, Kushal, Schmidt, Antonia, Weiser, Yan, Beyer, Anne, Schlangen, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024)
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
di: Beyer, Anne, et al.
Pubblicazione: (2024)
di: Beyer, Anne, et al.
Pubblicazione: (2024)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
di: Sadler, Philipp, et al.
Pubblicazione: (2024)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
di: Jordan, Jonathan, et al.
Pubblicazione: (2025)
di: Jordan, Jonathan, et al.
Pubblicazione: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2025)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2026)
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2026)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
di: Schlangen, David, et al.
Pubblicazione: (2025)
di: Schlangen, David, et al.
Pubblicazione: (2025)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
di: Thakkar, Gaurish, et al.
Pubblicazione: (2024)
di: Thakkar, Gaurish, et al.
Pubblicazione: (2024)
TurkicNLP: An NLP Toolkit for Turkic Languages
di: Hakimov, Sherzod
Pubblicazione: (2026)
di: Hakimov, Sherzod
Pubblicazione: (2026)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
di: Hakimov, Sherzod, et al.
Pubblicazione: (2023)
di: Hakimov, Sherzod, et al.
Pubblicazione: (2023)
LLMs as Function Approximators: Terminology, Taxonomy, and Questions for Evaluation
di: Schlangen, David
Pubblicazione: (2024)
di: Schlangen, David
Pubblicazione: (2024)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
di: Kang, Weitai, et al.
Pubblicazione: (2025)
di: Kang, Weitai, et al.
Pubblicazione: (2025)
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
di: Li, Meng, et al.
Pubblicazione: (2025)
di: Li, Meng, et al.
Pubblicazione: (2025)
Playpen: An Environment for Exploring Learning Through Conversational Interaction
di: Horst, Nicola, et al.
Pubblicazione: (2025)
di: Horst, Nicola, et al.
Pubblicazione: (2025)
Free-text Rationale Generation under Readability Level Control
di: Hsu, Yi-Sheng, et al.
Pubblicazione: (2024)
di: Hsu, Yi-Sheng, et al.
Pubblicazione: (2024)
Prior Lessons of Incremental Dialogue and Robot Action Management for the Age of Language Models
di: Kennington, Casey, et al.
Pubblicazione: (2025)
di: Kennington, Casey, et al.
Pubblicazione: (2025)
It Couldn't Help But Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised Learning
di: Madureira, Brielen, et al.
Pubblicazione: (2024)
di: Madureira, Brielen, et al.
Pubblicazione: (2024)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
di: Huo, Jiahao, et al.
Pubblicazione: (2025)
di: Huo, Jiahao, et al.
Pubblicazione: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
di: Chen, Jiaxing, et al.
Pubblicazione: (2024)
di: Chen, Jiaxing, et al.
Pubblicazione: (2024)
Plan-Grounded Large Language Models for Dual Goal Conversational Settings
di: Glória-Silva, Diogo, et al.
Pubblicazione: (2024)
di: Glória-Silva, Diogo, et al.
Pubblicazione: (2024)
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
di: Yu, Pengfei, et al.
Pubblicazione: (2025)
di: Yu, Pengfei, et al.
Pubblicazione: (2025)
Could the Road to Grounded, Neuro-symbolic AI be Paved with Words-as-Classifiers?
di: Kennington, Casey, et al.
Pubblicazione: (2025)
di: Kennington, Casey, et al.
Pubblicazione: (2025)
Multimodal Policy Internalization for Conversational Agents
di: Wang, Zhenhailong, et al.
Pubblicazione: (2025)
di: Wang, Zhenhailong, et al.
Pubblicazione: (2025)
Credence Calibration Game? Calibrating Large Language Models through Structured Play
di: Fang, Ke, et al.
Pubblicazione: (2025)
di: Fang, Ke, et al.
Pubblicazione: (2025)
Towards Visual Text Grounding of Multimodal Large Language Model
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
di: Zeng, Gailun, et al.
Pubblicazione: (2025)
di: Zeng, Gailun, et al.
Pubblicazione: (2025)
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
di: Jia, Boyu, et al.
Pubblicazione: (2025)
di: Jia, Boyu, et al.
Pubblicazione: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
di: Wang, Haibo, et al.
Pubblicazione: (2024)
di: Wang, Haibo, et al.
Pubblicazione: (2024)
A Dialogue Game for Eliciting Balanced Collaboration
di: Jeknić, Isidora, et al.
Pubblicazione: (2024)
di: Jeknić, Isidora, et al.
Pubblicazione: (2024)
Model Composition for Multimodal Large Language Models
di: Chen, Chi, et al.
Pubblicazione: (2024)
di: Chen, Chi, et al.
Pubblicazione: (2024)
DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base
di: Mao, Song, et al.
Pubblicazione: (2025)
di: Mao, Song, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
di: Hakimov, Sherzod, et al.
Pubblicazione: (2026) -
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
di: Bhavsar, Nidhir, et al.
Pubblicazione: (2024) -
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
di: Hakimov, Sherzod, et al.
Pubblicazione: (2025) -
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
di: Kranti, Chalamalasetti, et al.
Pubblicazione: (2024) -
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
di: Beyer, Anne, et al.
Pubblicazione: (2024)