CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nguyen, Kiet A., Juvekar, Adheesh, Yu, Tianjiao, Wahed, Muntasir, Lourentzou, Ismini
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913775504850944
author Nguyen, Kiet A.
Juvekar, Adheesh
Yu, Tianjiao
Wahed, Muntasir
Lourentzou, Ismini
author_facet Nguyen, Kiet A.
Juvekar, Adheesh
Yu, Tianjiao
Wahed, Muntasir
Lourentzou, Ismini
contents Recent advances in Large Vision-Language Models (LVLMs) have enabled general-purpose vision tasks through visual instruction tuning. While existing LVLMs can generate segmentation masks from text prompts for single images, they struggle with segmentation-grounded reasoning across images, especially at finer granularities such as object parts. In this paper, we introduce the new task of part-focused semantic co-segmentation, which involves identifying and segmenting common objects, as well as common and unique object parts across images. To address this task, we present CALICO, the first LVLM designed for multi-image part-level reasoning segmentation. CALICO features two key components, a novel Correspondence Extraction Module that identifies semantic part-level correspondences, and Correspondence Adaptation Modules that embed this information into the LVLM to facilitate multi-image understanding in a parameter-efficient manner. To support training and evaluation, we curate MixedParts, a large-scale multi-image segmentation dataset containing $\sim$2.4M samples across $\sim$44K images spanning diverse object and part categories. Experimental results demonstrate that CALICO, with just 0.3% of its parameters finetuned, achieves strong performance on this challenging task.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
Nguyen, Kiet A.
Juvekar, Adheesh
Yu, Tianjiao
Wahed, Muntasir
Lourentzou, Ismini
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in Large Vision-Language Models (LVLMs) have enabled general-purpose vision tasks through visual instruction tuning. While existing LVLMs can generate segmentation masks from text prompts for single images, they struggle with segmentation-grounded reasoning across images, especially at finer granularities such as object parts. In this paper, we introduce the new task of part-focused semantic co-segmentation, which involves identifying and segmenting common objects, as well as common and unique object parts across images. To address this task, we present CALICO, the first LVLM designed for multi-image part-level reasoning segmentation. CALICO features two key components, a novel Correspondence Extraction Module that identifies semantic part-level correspondences, and Correspondence Adaptation Modules that embed this information into the LVLM to facilitate multi-image understanding in a parameter-efficient manner. To support training and evaluation, we curate MixedParts, a large-scale multi-image segmentation dataset containing $\sim$2.4M samples across $\sim$44K images spanning diverse object and part categories. Experimental results demonstrate that CALICO, with just 0.3% of its parameters finetuned, achieves strong performance on this challenging task.
title CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.19331