The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niu, Quanzhu, Gong, Dengxian, Chen, Shihao, Zhang, Tao, Zhou, Yikang, Yuan, Haobo, Qi, Lu, Li, Xiangtai, Ji, Shunping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911219219169280
author Niu, Quanzhu
Gong, Dengxian
Chen, Shihao
Zhang, Tao
Zhou, Yikang
Yuan, Haobo
Qi, Lu
Li, Xiangtai
Ji, Shunping
author_facet Niu, Quanzhu
Gong, Dengxian
Chen, Shihao
Zhang, Tao
Zhou, Yikang
Yuan, Haobo
Qi, Lu
Li, Xiangtai
Ji, Shunping
contents Referring video object segmentation (RVOS) requires segmenting and tracking objects in videos conditioned on natural-language expressions, demanding fine-grained understanding of both appearance and motion. Building on Sa2VA, which couples a Multi-modal Large Language Model (MLLM) with the video segmentation model SAM2, we identify two key bottlenecks that limit segmentation performance: sparse frame sampling and reliance on a single [SEG] token for an entire video. We propose Segmentation Augmented and Selective Averaged Sa2VA (SaSaSa2VA) to address these issues. On the 7th LSVOS Challenge (RVOS track), SaSaSa2VA achieves a $\mathcal{J\&F}$ of 67.45, ranking first and surpassing the runner-up by 2.80 points. This result and ablation studies demonstrate that efficient segmentation augmentation and test-time ensembling substantially enhance grounded MLLMs for RVOS. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16972
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
Niu, Quanzhu
Gong, Dengxian
Chen, Shihao
Zhang, Tao
Zhou, Yikang
Yuan, Haobo
Qi, Lu
Li, Xiangtai
Ji, Shunping
Computer Vision and Pattern Recognition
Artificial Intelligence
Referring video object segmentation (RVOS) requires segmenting and tracking objects in videos conditioned on natural-language expressions, demanding fine-grained understanding of both appearance and motion. Building on Sa2VA, which couples a Multi-modal Large Language Model (MLLM) with the video segmentation model SAM2, we identify two key bottlenecks that limit segmentation performance: sparse frame sampling and reliance on a single [SEG] token for an entire video. We propose Segmentation Augmented and Selective Averaged Sa2VA (SaSaSa2VA) to address these issues. On the 7th LSVOS Challenge (RVOS track), SaSaSa2VA achieves a $\mathcal{J\&F}$ of 67.45, ranking first and surpassing the runner-up by 2.80 points. This result and ablation studies demonstrate that efficient segmentation augmentation and test-time ensembling substantially enhance grounded MLLMs for RVOS. The code is released in Sa2VA repository: https://github.com/bytedance/Sa2VA.
title The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.16972