Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pantazopoulos, Georgios, Suglia, Alessandro, Lemon, Oliver, Eshghi, Arash
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914763937677312
author Pantazopoulos, Georgios
Suglia, Alessandro
Lemon, Oliver
Eshghi, Arash
author_facet Pantazopoulos, Georgios
Suglia, Alessandro
Lemon, Oliver
Eshghi, Arash
contents An effective method for combining frozen large language models (LLM) and visual encoders involves a resampler module that creates a `visual prompt' which is provided to the LLM, along with the textual prompt. While this approach has enabled impressive performance across many coarse-grained tasks like image captioning and visual question answering, more fine-grained tasks that require spatial understanding have not been thoroughly examined. In this paper, we use \textit{diagnostic classifiers} to measure the extent to which the visual prompt produced by the resampler encodes spatial information. Our results show that this information is largely absent from the resampler output when kept frozen during training of the classifiers. However, when the resampler and classifier are trained jointly, we observe a significant performance boost. This shows that the compression achieved by the resamplers can in principle encode the requisite spatial information, but that more object-aware objectives are needed at the pretraining stage to facilitate this capability
format Preprint
id arxiv_https___arxiv_org_abs_2404_13594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
Pantazopoulos, Georgios
Suglia, Alessandro
Lemon, Oliver
Eshghi, Arash
Computer Vision and Pattern Recognition
Artificial Intelligence
An effective method for combining frozen large language models (LLM) and visual encoders involves a resampler module that creates a `visual prompt' which is provided to the LLM, along with the textual prompt. While this approach has enabled impressive performance across many coarse-grained tasks like image captioning and visual question answering, more fine-grained tasks that require spatial understanding have not been thoroughly examined. In this paper, we use \textit{diagnostic classifiers} to measure the extent to which the visual prompt produced by the resampler encodes spatial information. Our results show that this information is largely absent from the resampler output when kept frozen during training of the classifiers. However, when the resampler and classifier are trained jointly, we observe a significant performance boost. This shows that the compression achieved by the resamplers can in principle encode the requisite spatial information, but that more object-aware objectives are needed at the pretraining stage to facilitate this capability
title Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2404.13594