Evaluating SAM2 for Video Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ariff, Syed Hesham Syed, Liu, Yun, Sun, Guolei, Yang, Jing, Ding, Henghui, Geng, Xue, Jiang, Xudong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910004061143040
author Ariff, Syed Hesham Syed
Liu, Yun
Sun, Guolei
Yang, Jing
Ding, Henghui
Geng, Xue
Jiang, Xudong
author_facet Ariff, Syed Hesham Syed
Liu, Yun
Sun, Guolei
Yang, Jing
Ding, Henghui
Geng, Xue
Jiang, Xudong
contents The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them temporally through memory blocks. While SAM2 excels in video object segmentation by providing dense segmentation masks based on prompts, extending it to dense Video Semantic Segmentation (VSS) poses challenges due to the need for spatial accuracy, temporal consistency, and the ability to track multiple objects with complex boundaries and varying scales. This paper explores the extension of SAM2 for VSS, focusing on two primary approaches and highlighting firsthand observations and common challenges faced during this process. The first approach involves using SAM2 to extract unique objects as masks from a given image, with a segmentation network employed in parallel to generate and refine initial predictions. The second approach utilizes the predicted masks to extract unique feature vectors, which are then fed into a simple network for classification. The resulting classifications and masks are subsequently combined to produce the final segmentation. Our experiments suggest that leveraging SAM2 enhances overall performance in VSS, primarily due to its precise predictions of object boundaries.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01774
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating SAM2 for Video Semantic Segmentation
Ariff, Syed Hesham Syed
Liu, Yun
Sun, Guolei
Yang, Jing
Ding, Henghui
Geng, Xue
Jiang, Xudong
Computer Vision and Pattern Recognition
The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them temporally through memory blocks. While SAM2 excels in video object segmentation by providing dense segmentation masks based on prompts, extending it to dense Video Semantic Segmentation (VSS) poses challenges due to the need for spatial accuracy, temporal consistency, and the ability to track multiple objects with complex boundaries and varying scales. This paper explores the extension of SAM2 for VSS, focusing on two primary approaches and highlighting firsthand observations and common challenges faced during this process. The first approach involves using SAM2 to extract unique objects as masks from a given image, with a segmentation network employed in parallel to generate and refine initial predictions. The second approach utilizes the predicted masks to extract unique feature vectors, which are then fed into a simple network for classification. The resulting classifications and masks are subsequently combined to produce the final segmentation. Our experiments suggest that leveraging SAM2 enhances overall performance in VSS, primarily due to its precise predictions of object boundaries.
title Evaluating SAM2 for Video Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.01774