PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dorkenwald, Michael, Barazani, Nimrod, Snoek, Cees G. M., Asano, Yuki M.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916123761442816
author Dorkenwald, Michael
Barazani, Nimrod
Snoek, Cees G. M.
Asano, Yuki M.
author_facet Dorkenwald, Michael
Barazani, Nimrod
Snoek, Cees G. M.
Asano, Yuki M.
contents Vision-Language Models (VLMs), such as Flamingo and GPT-4V, have shown immense potential by integrating large language models with vision systems. Nevertheless, these models face challenges in the fundamental computer vision task of object localisation, due to their training on multimodal data containing mostly captions without explicit spatial grounding. While it is possible to construct custom, supervised training pipelines with bounding box annotations that integrate with VLMs, these result in specialized and hard-to-scale models. In this paper, we aim to explore the limits of caption-based VLMs and instead propose to tackle the challenge in a simpler manner by i) keeping the weights of a caption-based VLM frozen and ii) not using any supervised detection data. To this end, we introduce an input-agnostic Positional Insert (PIN), a learnable spatial prompt, containing a minimal set of parameters that are slid inside the frozen VLM, unlocking object localisation capabilities. Our PIN module is trained with a simple next-token prediction task on synthetic data without requiring the introduction of new output heads. Our experiments demonstrate strong zero-shot localisation performances on a variety of images, including Pascal VOC, COCO, LVIS, and diverse images like paintings or cartoons.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08657
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
Dorkenwald, Michael
Barazani, Nimrod
Snoek, Cees G. M.
Asano, Yuki M.
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs), such as Flamingo and GPT-4V, have shown immense potential by integrating large language models with vision systems. Nevertheless, these models face challenges in the fundamental computer vision task of object localisation, due to their training on multimodal data containing mostly captions without explicit spatial grounding. While it is possible to construct custom, supervised training pipelines with bounding box annotations that integrate with VLMs, these result in specialized and hard-to-scale models. In this paper, we aim to explore the limits of caption-based VLMs and instead propose to tackle the challenge in a simpler manner by i) keeping the weights of a caption-based VLM frozen and ii) not using any supervised detection data. To this end, we introduce an input-agnostic Positional Insert (PIN), a learnable spatial prompt, containing a minimal set of parameters that are slid inside the frozen VLM, unlocking object localisation capabilities. Our PIN module is trained with a simple next-token prediction task on synthetic data without requiring the introduction of new output heads. Our experiments demonstrate strong zero-shot localisation performances on a variety of images, including Pascal VOC, COCO, LVIS, and diverse images like paintings or cartoons.
title PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.08657