Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
Fuente:
arXiv
Guardado en:
| Autores principales: | Kranti, Chalamalasetti, Hakimov, Sherzod, Schlangen, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
por: Kranti, Chalamalasetti, et al.
Publicado: (2024)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
por: Schlangen, David, et al.
Publicado: (2025)
por: Schlangen, David, et al.
Publicado: (2025)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
por: Beyer, Anne, et al.
Publicado: (2024)
por: Beyer, Anne, et al.
Publicado: (2024)
Sharing the Cost of Success: A Game for Evaluating and Learning Collaborative Multi-Agent Instruction Giving and Following Policies
por: Sadler, Philipp, et al.
Publicado: (2024)
por: Sadler, Philipp, et al.
Publicado: (2024)
Learning Communication Policies for Different Follower Behaviors in a Collaborative Reference Game
por: Sadler, Philipp, et al.
Publicado: (2024)
por: Sadler, Philipp, et al.
Publicado: (2024)
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
por: Hakimov, Sherzod, et al.
Publicado: (2026)
por: Hakimov, Sherzod, et al.
Publicado: (2026)
Ad-hoc Concept Forming in the Game Codenames as a Means for Evaluating Large Language Models
por: Hakimov, Sherzod, et al.
Publicado: (2025)
por: Hakimov, Sherzod, et al.
Publicado: (2025)
Plant in Cupboard, Orange on Rably, Inat Aphone. Benchmarking Incremental Learning of Situation and Language Model using a Text-Simulated Situated Environment
por: Jordan, Jonathan, et al.
Publicado: (2025)
por: Jordan, Jonathan, et al.
Publicado: (2025)
How Many Parameters Does it Take to Change a Light Bulb? Evaluating Performance in Self-Play of Conversational Games as a Function of Model Characteristics
por: Bhavsar, Nidhir, et al.
Publicado: (2024)
por: Bhavsar, Nidhir, et al.
Publicado: (2024)
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
TurkicNLP: An NLP Toolkit for Turkic Languages
por: Hakimov, Sherzod
Publicado: (2026)
por: Hakimov, Sherzod
Publicado: (2026)
The Price of Thought: A Multilingual Analysis of Reasoning, Performance, and Cost of Negotiation in Large Language Models
por: Hakimov, Sherzod, et al.
Publicado: (2025)
por: Hakimov, Sherzod, et al.
Publicado: (2025)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
por: Kranti, Chalamalasetti
Publicado: (2025)
por: Kranti, Chalamalasetti
Publicado: (2025)
A Dialogue Game for Eliciting Balanced Collaboration
por: Jeknić, Isidora, et al.
Publicado: (2024)
por: Jeknić, Isidora, et al.
Publicado: (2024)
Unveiling Global Narratives: A Multilingual Twitter Dataset of News Media on the Russo-Ukrainian Conflict
por: Hakimov, Sherzod, et al.
Publicado: (2023)
por: Hakimov, Sherzod, et al.
Publicado: (2023)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
por: Chen, Boyuan, et al.
Publicado: (2024)
por: Chen, Boyuan, et al.
Publicado: (2024)
Using Game Play to Investigate Multimodal and Conversational Grounding in Large Multimodal Models
por: Hakimov, Sherzod, et al.
Publicado: (2024)
por: Hakimov, Sherzod, et al.
Publicado: (2024)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
por: Wang, Kangrui, et al.
Publicado: (2025)
por: Wang, Kangrui, et al.
Publicado: (2025)
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
por: Thakkar, Gaurish, et al.
Publicado: (2024)
por: Thakkar, Gaurish, et al.
Publicado: (2024)
Test Set Quality in Multilingual LLM Evaluation
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
por: Kranti, Chalamalasetti, et al.
Publicado: (2025)
Can Visual Dialogue Models Do Scorekeeping? Exploring How Dialogue Representations Incrementally Encode Shared Knowledge
por: Madureira, Brielen, et al.
Publicado: (2022)
por: Madureira, Brielen, et al.
Publicado: (2022)
A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment
por: Inoue, Koji, et al.
Publicado: (2025)
por: Inoue, Koji, et al.
Publicado: (2025)
Coherence-Driven Multimodal Safety Dialogue with Active Learning for Embodied Agents
por: Hassan, Sabit, et al.
Publicado: (2024)
por: Hassan, Sabit, et al.
Publicado: (2024)
Free-text Rationale Generation under Readability Level Control
por: Hsu, Yi-Sheng, et al.
Publicado: (2024)
por: Hsu, Yi-Sheng, et al.
Publicado: (2024)
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues
por: Kuo, Martin, et al.
Publicado: (2025)
por: Kuo, Martin, et al.
Publicado: (2025)
SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue
por: Dai, Yuqin, et al.
Publicado: (2026)
por: Dai, Yuqin, et al.
Publicado: (2026)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
por: Zhang, Yang, et al.
Publicado: (2024)
por: Zhang, Yang, et al.
Publicado: (2024)
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
por: Badola, Kartikeya, et al.
Publicado: (2025)
por: Badola, Kartikeya, et al.
Publicado: (2025)
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
por: Yuan, Puzhen, et al.
Publicado: (2025)
por: Yuan, Puzhen, et al.
Publicado: (2025)
PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
por: Dao, Alan, et al.
Publicado: (2025)
por: Dao, Alan, et al.
Publicado: (2025)
When Only Time Will Tell: Interpreting How Transformers Process Local Ambiguities Through the Lens of Restart-Incrementality
por: Madureira, Brielen, et al.
Publicado: (2024)
por: Madureira, Brielen, et al.
Publicado: (2024)
Leveraging Adaptive Group Negotiation for Heterogeneous Multi-Robot Collaboration with Large Language Models
por: Song, Siqi, et al.
Publicado: (2025)
por: Song, Siqi, et al.
Publicado: (2025)
Autonomous Frontier-Based Exploration with VLM Guidance
por: Aitha, Aarush, et al.
Publicado: (2026)
por: Aitha, Aarush, et al.
Publicado: (2026)
Dialogue with Robots: Proposals for Broadening Participation and Research in the SLIVAR Community
por: Kennington, Casey, et al.
Publicado: (2024)
por: Kennington, Casey, et al.
Publicado: (2024)
Applying General Turn-taking Models to Conversational Human-Robot Interaction
por: Skantze, Gabriel, et al.
Publicado: (2025)
por: Skantze, Gabriel, et al.
Publicado: (2025)
Multi-Agent Consensus Seeking via Large Language Models
por: Chen, Huaben, et al.
Publicado: (2023)
por: Chen, Huaben, et al.
Publicado: (2023)
Ejemplares similares
-
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
por: Kranti, Chalamalasetti, et al.
Publicado: (2024) -
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
por: Kranti, Chalamalasetti, et al.
Publicado: (2025) -
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
por: Kranti, Chalamalasetti, et al.
Publicado: (2025) -
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
por: Kranti, Chalamalasetti, et al.
Publicado: (2024) -
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
por: Schlangen, David, et al.
Publicado: (2025)