Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pätzold, Bastian, Nogga, Jan, Behnke, Sven
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908601026609152
author Pätzold, Bastian
Nogga, Jan
Behnke, Sven
author_facet Pätzold, Bastian
Nogga, Jan
Behnke, Sven
contents Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary object detection (OVD), instance segmentation, and tracking leverages their strengths while mitigating these drawbacks. We utilize VLM-generated structured descriptions to identify visible object instances, collect application-relevant attributes, and inform an open-vocabulary detector to extract corresponding bounding boxes that are passed to a video segmentation model providing segmentation masks and tracking. Once initialized, this model directly extracts segmentation masks, processing image streams in real time with minimal computational overhead. Tracks can be updated online as needed by generating new structured descriptions and detections. This combines the descriptive power of VLMs with the grounding capability of OVD and the pixel-level understanding and speed of video segmentation. Our evaluation across datasets and robotics platforms demonstrates the broad applicability of this approach, showcasing its ability to extract task-specific attributes from non-standard objects in dynamic environments. Code, data, videos, and benchmarks are available at https://vlm-gist.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2503_16538
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
Pätzold, Bastian
Nogga, Jan
Behnke, Sven
Computer Vision and Pattern Recognition
Robotics
Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary object detection (OVD), instance segmentation, and tracking leverages their strengths while mitigating these drawbacks. We utilize VLM-generated structured descriptions to identify visible object instances, collect application-relevant attributes, and inform an open-vocabulary detector to extract corresponding bounding boxes that are passed to a video segmentation model providing segmentation masks and tracking. Once initialized, this model directly extracts segmentation masks, processing image streams in real time with minimal computational overhead. Tracks can be updated online as needed by generating new structured descriptions and detections. This combines the descriptive power of VLMs with the grounding capability of OVD and the pixel-level understanding and speed of video segmentation. Our evaluation across datasets and robotics platforms demonstrates the broad applicability of this approach, showcasing its ability to extract task-specific attributes from non-standard objects in dynamic environments. Code, data, videos, and benchmarks are available at https://vlm-gist.github.io
title Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.16538