PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914175247187968 |
|---|---|
| author | Wahed, Muntasir Nguyen, Kiet A. Juvekar, Adheesh Sunil Li, Xinzhuo Zhou, Xiaona Shah, Vedant Yu, Tianjiao Yanardag, Pinar Lourentzou, Ismini |
| author_facet | Wahed, Muntasir Nguyen, Kiet A. Juvekar, Adheesh Sunil Li, Xinzhuo Zhou, Xiaona Shah, Vedant Yu, Tianjiao Yanardag, Pinar Lourentzou, Ismini |
| contents | Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning alongside PRIMA, an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE, a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4SEG, a new multi-image reasoning segmentation benchmark consisting of $\sim$744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with $7.83\%$ and $11.25\%$ improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_15209 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation Wahed, Muntasir Nguyen, Kiet A. Juvekar, Adheesh Sunil Li, Xinzhuo Zhou, Xiaona Shah, Vedant Yu, Tianjiao Yanardag, Pinar Lourentzou, Ismini Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Despite significant advancements in Large Vision-Language Models (LVLMs)' capabilities, existing pixel-grounding models operate in single-image settings, limiting their ability to perform detailed, fine-grained comparisons across multiple images. Conversely, current multi-image understanding models lack pixel-level grounding. Our work addresses this gap by introducing the task of multi-image pixel-grounded reasoning alongside PRIMA, an LVLM that integrates pixel-level grounding with robust multi-image reasoning to produce contextually rich, pixel-grounded explanations. Central to PRIMA is SQuARE, a vision module that injects cross-image relational context into compact query-based visual tokens before fusing them with the language backbone. To support training and evaluation, we curate M4SEG, a new multi-image reasoning segmentation benchmark consisting of $\sim$744K question-answer pairs that require fine-grained visual understanding across multiple images. PRIMA outperforms state-of-the-art baselines with $7.83\%$ and $11.25\%$ improvements in Recall and S-IoU, respectively. Ablation studies further demonstrate the effectiveness of the proposed SQuARE module in capturing cross-image relationships. |
| title | PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2412.15209 |