Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Haobo, Li, Xiangtai, Zhang, Tao, Sun, Yueyi, Huang, Zilong, Xu, Shilin, Ji, Shunping, Tong, Yunhai, Qi, Lu, Feng, Jiashi, Yang, Ming-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911247186788352
author Yuan, Haobo
Li, Xiangtai
Zhang, Tao
Sun, Yueyi
Huang, Zilong
Xu, Shilin
Ji, Shunping
Tong, Yunhai
Qi, Lu
Feng, Jiashi
Yang, Ming-Hsuan
author_facet Yuan, Haobo
Li, Xiangtai
Zhang, Tao
Sun, Yueyi
Huang, Zilong
Xu, Shilin
Ji, Shunping
Tong, Yunhai
Qi, Lu
Feng, Jiashi
Yang, Ming-Hsuan
contents This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation, with minimal one-shot instruction tuning. Sa2VA combines SAM-2, a foundation video segmentation model, with MLLM, the advanced vision-language model, and unifies text, image, and video into a shared LLM token space. Using the LLM, Sa2VA generates instruction tokens that guide SAM-2 in producing precise masks, enabling a grounded, multi-modal understanding of both static and dynamic visual content. Additionally, we introduce Ref-SAV, an auto-labeled dataset containing over 72k object expressions in complex video scenes, designed to boost model performance. We also manually validate 2k video objects in the Ref-SAV datasets to benchmark referring video object segmentation in complex environments. Experiments show that Sa2VA achieves strong performance across multiple tasks, particularly in referring video object segmentation, highlighting its potential for complex real-world applications. In addition, Sa2VA can be easily extended into various VLMs, including Qwen-VL and Intern-VL, which can be updated with rapid process in current open-sourced VLMs. Code and models have been provided to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Yuan, Haobo
Li, Xiangtai
Zhang, Tao
Sun, Yueyi
Huang, Zilong
Xu, Shilin
Ji, Shunping
Tong, Yunhai
Qi, Lu
Feng, Jiashi
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-modal large language models, which are often limited to specific modalities and tasks, Sa2VA supports a wide range of image and video tasks, including referring segmentation and conversation, with minimal one-shot instruction tuning. Sa2VA combines SAM-2, a foundation video segmentation model, with MLLM, the advanced vision-language model, and unifies text, image, and video into a shared LLM token space. Using the LLM, Sa2VA generates instruction tokens that guide SAM-2 in producing precise masks, enabling a grounded, multi-modal understanding of both static and dynamic visual content. Additionally, we introduce Ref-SAV, an auto-labeled dataset containing over 72k object expressions in complex video scenes, designed to boost model performance. We also manually validate 2k video objects in the Ref-SAV datasets to benchmark referring video object segmentation in complex environments. Experiments show that Sa2VA achieves strong performance across multiple tasks, particularly in referring video object segmentation, highlighting its potential for complex real-world applications. In addition, Sa2VA can be easily extended into various VLMs, including Qwen-VL and Intern-VL, which can be updated with rapid process in current open-sourced VLMs. Code and models have been provided to the community.
title Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.04001