Grounding Language in Multi-Perspective Referential Communication

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tang, Zineng, Mao, Lingjun, Suhr, Alane
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913533792354304
author Tang, Zineng
Mao, Lingjun
Suhr, Alane
author_facet Tang, Zineng
Mao, Lingjun
Suhr, Alane
contents We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspective, which may be different from their own, to both produce and understand references to objects in a scene and the spatial relations between them. We collect a dataset of 2,970 human-written referring expressions, each paired with human comprehension judgments, and evaluate the performance of automated models as speakers and listeners paired with human partners, finding that model performance in both reference generation and comprehension lags behind that of pairs of human agents. Finally, we experiment training an open-weight speaker model with evidence of communicative success when paired with a listener, resulting in an improvement from 58.9 to 69.3% in communicative success and even outperforming the strongest proprietary model.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03959
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grounding Language in Multi-Perspective Referential Communication
Tang, Zineng
Mao, Lingjun
Suhr, Alane
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphics
We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspective, which may be different from their own, to both produce and understand references to objects in a scene and the spatial relations between them. We collect a dataset of 2,970 human-written referring expressions, each paired with human comprehension judgments, and evaluate the performance of automated models as speakers and listeners paired with human partners, finding that model performance in both reference generation and comprehension lags behind that of pairs of human agents. Finally, we experiment training an open-weight speaker model with evidence of communicative success when paired with a listener, resulting in an improvement from 58.9 to 69.3% in communicative success and even outperforming the strongest proprietary model.
title Grounding Language in Multi-Perspective Referential Communication
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2410.03959