SAMSON: 3rd Place Solution of LSVOS 2025 VOS Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yujie, Zhang, Hongyang, Liu, Zhihui, Ruan, Shihai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915506508791808
author Xie, Yujie
Zhang, Hongyang
Liu, Zhihui
Ruan, Shihai
author_facet Xie, Yujie
Zhang, Hongyang
Liu, Zhihui
Ruan, Shihai
contents Large-scale Video Object Segmentation (LSVOS) addresses the challenge of accurately tracking and segmenting objects in long video sequences, where difficulties stem from object reappearance, small-scale targets, heavy occlusions, and crowded scenes. Existing approaches predominantly adopt SAM2-based frameworks with various memory mechanisms for complex video mask generation. In this report, we proposed Segment Anything with Memory Strengthened Object Navigation (SAMSON), the 3rd place solution in the MOSE track of ICCV 2025, which integrates the strengths of stateof-the-art VOS models into an effective paradigm. To handle visually similar instances and long-term object disappearance in MOSE, we incorporate a long-term memorymodule for reliable object re-identification. Additionly, we adopt SAM2Long as a post-processing strategy to reduce error accumulation and enhance segmentation stability in long video sequences. Our method achieved a final performance of 0.8427 in terms of J &F in the test-set leaderboard.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAMSON: 3rd Place Solution of LSVOS 2025 VOS Challenge
Xie, Yujie
Zhang, Hongyang
Liu, Zhihui
Ruan, Shihai
Computer Vision and Pattern Recognition
Large-scale Video Object Segmentation (LSVOS) addresses the challenge of accurately tracking and segmenting objects in long video sequences, where difficulties stem from object reappearance, small-scale targets, heavy occlusions, and crowded scenes. Existing approaches predominantly adopt SAM2-based frameworks with various memory mechanisms for complex video mask generation. In this report, we proposed Segment Anything with Memory Strengthened Object Navigation (SAMSON), the 3rd place solution in the MOSE track of ICCV 2025, which integrates the strengths of stateof-the-art VOS models into an effective paradigm. To handle visually similar instances and long-term object disappearance in MOSE, we incorporate a long-term memorymodule for reliable object re-identification. Additionly, we adopt SAM2Long as a post-processing strategy to reduce error accumulation and enhance segmentation stability in long video sequences. Our method achieved a final performance of 0.8427 in terms of J &F in the test-set leaderboard.
title SAMSON: 3rd Place Solution of LSVOS 2025 VOS Challenge
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17500