Teaching VLMs to Localize Specific Objects from In-context Examples

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Doveh, Sivan, Shabtay, Nimrod, Lin, Wei, Schwartz, Eli, Kuehne, Hilde, Giryes, Raja, Feris, Rogerio, Karlinsky, Leonid, Glass, James, Arbelle, Assaf, Ullman, Shimon, Mirza, M. Jehanzeb
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913733389844480
author Doveh, Sivan
Shabtay, Nimrod
Lin, Wei
Schwartz, Eli
Kuehne, Hilde
Giryes, Raja
Feris, Rogerio
Karlinsky, Leonid
Glass, James
Arbelle, Assaf
Ullman, Shimon
Mirza, M. Jehanzeb
author_facet Doveh, Sivan
Shabtay, Nimrod
Lin, Wei
Schwartz, Eli
Kuehne, Hilde
Giryes, Raja
Feris, Rogerio
Karlinsky, Leonid
Glass, James
Arbelle, Assaf
Ullman, Shimon
Mirza, M. Jehanzeb
contents Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these advances, we find that present-day VLMs (including the proprietary GPT-4o) lack a fundamental cognitive ability: learning to localize specific objects in a scene by taking into account the context. In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with localizing the same object type in a query image. Personalized localization can be particularly important in cases of ambiguity of several related objects that can respond to a text or an object that is hard to describe with words. To provoke personalized localization abilities in models, we present a data-centric solution that fine-tunes them using carefully curated data from video object tracking datasets. By leveraging sequences of frames tracking the same object across multiple shots, we simulate instruction-tuning dialogues that promote context awareness. To reinforce this, we introduce a novel regularization technique that replaces object labels with pseudo-names, ensuring the model relies on visual context rather than prior knowledge. Our method significantly enhances the few-shot localization performance of recent VLMs ranging from 7B to 72B in size, without sacrificing generalization, as demonstrated on several benchmarks tailored towards evaluating personalized localization abilities. This work is the first to explore and benchmark personalized few-shot localization for VLMs -- exposing critical weaknesses in present-day VLMs, and laying a foundation for future research in context-driven vision-language applications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_13317
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Teaching VLMs to Localize Specific Objects from In-context Examples
Doveh, Sivan
Shabtay, Nimrod
Lin, Wei
Schwartz, Eli
Kuehne, Hilde
Giryes, Raja
Feris, Rogerio
Karlinsky, Leonid
Glass, James
Arbelle, Assaf
Ullman, Shimon
Mirza, M. Jehanzeb
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these advances, we find that present-day VLMs (including the proprietary GPT-4o) lack a fundamental cognitive ability: learning to localize specific objects in a scene by taking into account the context. In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with localizing the same object type in a query image. Personalized localization can be particularly important in cases of ambiguity of several related objects that can respond to a text or an object that is hard to describe with words. To provoke personalized localization abilities in models, we present a data-centric solution that fine-tunes them using carefully curated data from video object tracking datasets. By leveraging sequences of frames tracking the same object across multiple shots, we simulate instruction-tuning dialogues that promote context awareness. To reinforce this, we introduce a novel regularization technique that replaces object labels with pseudo-names, ensuring the model relies on visual context rather than prior knowledge. Our method significantly enhances the few-shot localization performance of recent VLMs ranging from 7B to 72B in size, without sacrificing generalization, as demonstrated on several benchmarks tailored towards evaluating personalized localization abilities. This work is the first to explore and benchmark personalized few-shot localization for VLMs -- exposing critical weaknesses in present-day VLMs, and laying a foundation for future research in context-driven vision-language applications.
title Teaching VLMs to Localize Specific Objects from In-context Examples
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.13317