Visual Textualization for Image Prompted Object Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yongjian, Zhou, Yang, Saiyin, Jiya, Wei, Bingzheng, Xu, Yan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909667147382784
author Wu, Yongjian
Zhou, Yang
Saiyin, Jiya
Wei, Bingzheng
Xu, Yan
author_facet Wu, Yongjian
Zhou, Yang
Saiyin, Jiya
Wei, Bingzheng
Xu, Yan
contents We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual exemplars into the text feature space to enhance Object-level Vision-Language Models' (OVLMs) capability in detecting rare categories that are difficult to describe textually and nearly absent from their pre-training data, while preserving their pre-trained object-text alignment. Specifically, VisTex-OVLM leverages multi-scale textualizing blocks and a multi-stage fusion strategy to integrate visual information from visual exemplars, generating textualized visual tokens that effectively guide OVLMs alongside text prompts. Unlike previous methods, our method maintains the original architecture of OVLM, maintaining its generalization capabilities while enhancing performance in few-shot settings. VisTex-OVLM demonstrates superior performance across open-set datasets which have minimal overlap with OVLM's pre-training data and achieves state-of-the-art results on few-shot benchmarks PASCAL VOC and MSCOCO. The code will be released at https://github.com/WitGotFlg/VisTex-OVLM.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23785
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Textualization for Image Prompted Object Detection
Wu, Yongjian
Zhou, Yang
Saiyin, Jiya
Wei, Bingzheng
Xu, Yan
Computer Vision and Pattern Recognition
We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization -- a process that projects a few visual exemplars into the text feature space to enhance Object-level Vision-Language Models' (OVLMs) capability in detecting rare categories that are difficult to describe textually and nearly absent from their pre-training data, while preserving their pre-trained object-text alignment. Specifically, VisTex-OVLM leverages multi-scale textualizing blocks and a multi-stage fusion strategy to integrate visual information from visual exemplars, generating textualized visual tokens that effectively guide OVLMs alongside text prompts. Unlike previous methods, our method maintains the original architecture of OVLM, maintaining its generalization capabilities while enhancing performance in few-shot settings. VisTex-OVLM demonstrates superior performance across open-set datasets which have minimal overlap with OVLM's pre-training data and achieves state-of-the-art results on few-shot benchmarks PASCAL VOC and MSCOCO. The code will be released at https://github.com/WitGotFlg/VisTex-OVLM.
title Visual Textualization for Image Prompted Object Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23785