Measuring How (Not Just Whether) VLMs Build Common Ground

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Imai, Saki, İnan, Mert, Sicilia, Anthony, Alikhani, Malihe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911138186264576
author Imai, Saki
İnan, Mert
Sicilia, Anthony
Alikhani, Malihe
author_facet Imai, Saki
İnan, Mert
Sicilia, Anthony
Alikhani, Malihe
contents Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gradually develop shared understanding through ongoing communication. We introduce a four-metric suite (grounding efficiency, content alignment, lexical adaptation, and human-likeness) to systematically evaluate VLM performance in interactive grounding contexts. We deploy the suite on 150 self-play sessions of interactive referential games between three proprietary VLMs and compare them with human dyads. All three models diverge from human patterns on at least three metrics, while GPT4o-mini is the closest overall. We find that (i) task success scores do not indicate successful grounding and (ii) high image-utterance alignment does not necessarily predict task success. Our metric suite and findings offer a framework for future research on VLM grounding.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Measuring How (Not Just Whether) VLMs Build Common Ground
Imai, Saki
İnan, Mert
Sicilia, Anthony
Alikhani, Malihe
Computation and Language
Artificial Intelligence
Large vision language models (VLMs) increasingly claim reasoning skills, yet current benchmarks evaluate them in single-turn or question answering settings. However, grounding is an interactive process in which people gradually develop shared understanding through ongoing communication. We introduce a four-metric suite (grounding efficiency, content alignment, lexical adaptation, and human-likeness) to systematically evaluate VLM performance in interactive grounding contexts. We deploy the suite on 150 self-play sessions of interactive referential games between three proprietary VLMs and compare them with human dyads. All three models diverge from human patterns on at least three metrics, while GPT4o-mini is the closest overall. We find that (i) task success scores do not indicate successful grounding and (ii) high image-utterance alignment does not necessarily predict task success. Our metric suite and findings offer a framework for future research on VLM grounding.
title Measuring How (Not Just Whether) VLMs Build Common Ground
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.03805