MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deichler, Anna, O'Regan, Jim, Dogan, Fethiye Irmak, Marcinek, Lubos, Klezovich, Anna, Leite, Iolanda, Beskow, Jonas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910242890055680
author Deichler, Anna
O'Regan, Jim
Dogan, Fethiye Irmak
Marcinek, Lubos
Klezovich, Anna
Leite, Iolanda
Beskow, Jonas
author_facet Deichler, Anna
O'Regan, Jim
Dogan, Fethiye Irmak
Marcinek, Lubos
Klezovich, Anna
Leite, Iolanda
Beskow, Jonas
contents Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous expressions in spontaneous, multi-turn dialogue. We address this gap by introducing (1) a benchmark for referential communication in dynamic 3D environments, built from 6.7 hours of egocentric VR interaction with synchronized speech, motion, gaze, and 3D scene geometry, and (2) a two-stage grounding pipeline that explicitly resolves conversational ambiguity before visual localization. The benchmark includes over 4,200 manually verified referring expressions spanning full, partitive, and pronominal types. Our contextual rewriting approach improves grounding performance by 11-22 percentage points on average, with a pure detector (GroundingDINO) reaching 56.7% on pronominals after rewriting, nearly double the best end-to-end baseline. Results demonstrate that decoupling linguistic reasoning from visual perception is more effective than end-to-end approaches for conversational grounding.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21796
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
Deichler, Anna
O'Regan, Jim
Dogan, Fethiye Irmak
Marcinek, Lubos
Klezovich, Anna
Leite, Iolanda
Beskow, Jonas
Computer Vision and Pattern Recognition
Computation and Language
I.2.7; I.2.10; H.5.2
Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to resolve ambiguous expressions in spontaneous, multi-turn dialogue. We address this gap by introducing (1) a benchmark for referential communication in dynamic 3D environments, built from 6.7 hours of egocentric VR interaction with synchronized speech, motion, gaze, and 3D scene geometry, and (2) a two-stage grounding pipeline that explicitly resolves conversational ambiguity before visual localization. The benchmark includes over 4,200 manually verified referring expressions spanning full, partitive, and pronominal types. Our contextual rewriting approach improves grounding performance by 11-22 percentage points on average, with a pure detector (GroundingDINO) reaching 56.7% on pronominals after rewriting, nearly double the best end-to-end baseline. Results demonstrate that decoupling linguistic reasoning from visual perception is more effective than end-to-end approaches for conversational grounding.
title MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
topic Computer Vision and Pattern Recognition
Computation and Language
I.2.7; I.2.10; H.5.2
url https://arxiv.org/abs/2605.21796