SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Dengxian, Niu, Quanzhu, Chen, Shihao, Wu, Yuanzheng, Zhou, Yikang, Zhang, Tao, Yuan, Haobo, Qi, Lu, Ji, Shunping
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908918978969600
author Gong, Dengxian
Niu, Quanzhu
Chen, Shihao
Wu, Yuanzheng
Zhou, Yikang
Zhang, Tao
Yuan, Haobo
Qi, Lu
Ji, Shunping
author_facet Gong, Dengxian
Niu, Quanzhu
Chen, Shihao
Wu, Yuanzheng
Zhou, Yikang
Zhang, Tao
Yuan, Haobo
Qi, Lu
Ji, Shunping
contents Referring video object segmentation (RVOS) commonly grounds targets in videos based on static textual cues. MeViS benchmark extends this by incorporating motion-centric expressions (referring & reasoning motion expressions) and introducing no-target queries. Extending SaSaSa2VA, where increased input frames and [SEG] tokens already strengthen the Sa2VA backbone, we adopt a simple yet effective target existence-aware verification mechanism, leading to Still Awesome SaSaSa2VA (SaSaSaSa2VA). Despite its simplicity, the method achieves a final score of 89.19 in the 5th PVUW Challenge (MeViS-Text Track), securing 2nd place. Both quantitative results and ablations suggest that this existence-aware verification strategy is sufficient to unlock strong performance on motion-centric referring tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27241
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
Gong, Dengxian
Niu, Quanzhu
Chen, Shihao
Wu, Yuanzheng
Zhou, Yikang
Zhang, Tao
Yuan, Haobo
Qi, Lu
Ji, Shunping
Computer Vision and Pattern Recognition
Referring video object segmentation (RVOS) commonly grounds targets in videos based on static textual cues. MeViS benchmark extends this by incorporating motion-centric expressions (referring & reasoning motion expressions) and introducing no-target queries. Extending SaSaSa2VA, where increased input frames and [SEG] tokens already strengthen the Sa2VA backbone, we adopt a simple yet effective target existence-aware verification mechanism, leading to Still Awesome SaSaSa2VA (SaSaSaSa2VA). Despite its simplicity, the method achieves a final score of 89.19 in the 5th PVUW Challenge (MeViS-Text Track), securing 2nd place. Both quantitative results and ablations suggest that this existence-aware verification strategy is sufficient to unlock strong performance on motion-centric referring tasks.
title SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27241