PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wahed, Muntasir, Nguyen, Kiet A., Juvekar, Adheesh Sunil, Li, Xinzhuo, Zhou, Xiaona, Shah, Vedant, Yu, Tianjiao, Yanardag, Pinar, Lourentzou, Ismini
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914175247187968
author Wahed, Muntasir
Nguyen, Kiet A.
Juvekar, Adheesh Sunil
Li, Xinzhuo
Zhou, Xiaona
Shah, Vedant
Yu, Tianjiao
Yanardag, Pinar
Lourentzou, Ismini
author_facet Wahed, Muntasir
Nguyen, Kiet A.
Juvekar, Adheesh Sunil
Li, Xinzhuo
Zhou, Xiaona
Shah, Vedant
Yu, Tianjiao
Yanardag, Pinar
Lourentzou, Ismini
contents Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning alongside PRIMA, an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE, a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4SEG, a new multi-image reasoning segmentation benchmark consisting of $\sim$744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with $7.83\%$ and $11.25\%$ improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15209
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
Wahed, Muntasir
Nguyen, Kiet A.
Juvekar, Adheesh Sunil
Li, Xinzhuo
Zhou, Xiaona
Shah, Vedant
Yu, Tianjiao
Yanardag, Pinar
Lourentzou, Ismini
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning alongside PRIMA, an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE, a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4SEG, a new multi-image reasoning segmentation benchmark consisting of $\sim$744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with $7.83\%$ and $11.25\%$ improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships.
title PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.15209