Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mitra, Chancharik, Anwar, Abrar, Corona, Rodolfo, Klein, Dan, Darrell, Trevor, Thomason, Jesse
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914747268464640
author Mitra, Chancharik
Anwar, Abrar
Corona, Rodolfo
Klein, Dan
Darrell, Trevor
Thomason, Jesse
author_facet Mitra, Chancharik
Anwar, Abrar
Corona, Rodolfo
Klein, Dan
Darrell, Trevor
Thomason, Jesse
contents When connecting objects and their language referents in an embodied 3D environment, it is important to note that: (1) an object can be better characterized by leveraging comparative information between itself and other objects, and (2) an object's appearance can vary with camera position. As such, we present the Multi-view Approach to Grounding in Context (MAGiC), which selects an object referent based on language that distinguishes between two similar objects. By pragmatically reasoning over both objects and across multiple views of those objects, MAGiC improves over the state-of-the-art model on the SNARE object reference task with a relative error reduction of 12.9\% (representing an absolute improvement of 2.7\%). Ablation studies show that reasoning jointly over object referent candidates and multiple views of each object both contribute to improved accuracy. Code: https://github.com/rcorona/magic_snare/
format Preprint
id arxiv_https___arxiv_org_abs_2311_06694
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
Mitra, Chancharik
Anwar, Abrar
Corona, Rodolfo
Klein, Dan
Darrell, Trevor
Thomason, Jesse
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
When connecting objects and their language referents in an embodied 3D environment, it is important to note that: (1) an object can be better characterized by leveraging comparative information between itself and other objects, and (2) an object's appearance can vary with camera position. As such, we present the Multi-view Approach to Grounding in Context (MAGiC), which selects an object referent based on language that distinguishes between two similar objects. By pragmatically reasoning over both objects and across multiple views of those objects, MAGiC improves over the state-of-the-art model on the SNARE object reference task with a relative error reduction of 12.9\% (representing an absolute improvement of 2.7\%). Ablation studies show that reasoning jointly over object referent candidates and multiple views of each object both contribute to improved accuracy. Code: https://github.com/rcorona/magic_snare/
title Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2311.06694