Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Nekrasov, Alexey, Athar, Ali, de Geus, Daan, Hermans, Alexander, Leibe, Bastian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908661326020608
author Nekrasov, Alexey
Athar, Ali
de Geus, Daan
Hermans, Alexander
Leibe, Bastian
author_facet Nekrasov, Alexey
Athar, Ali
de Geus, Daan
Hermans, Alexander
Leibe, Bastian
contents Sa2VA is a recent model for language-guided dense grounding in images and video that achieves state-of-the-art results on multiple segmentation benchmarks and that has become widely popular. However, we found that Sa2VA does not perform according to its full potential for referring video object segmentation tasks. We identify inconsistencies between training and inference procedures as the key factor holding it back. To mitigate this issue, we propose an improved version of Sa2VA, Sa2VA-i, that rectifies these issues and improves the results. In fact, Sa2VA-i sets a new state of the art for multiple video benchmarks and achieves improvements of up to +11.6 J&F on MeViS, +1.4 on Ref-YT-VOS, +3.3 on Ref-DAVIS and +4.1 on ReVOS using the same Sa2VA checkpoints. With our fixes, the Sa2VA-i-1B model even performs on par with the original Sa2VA-26B model on the MeViS benchmark. We hope that this work will show the importance of seemingly trivial implementation details and that it will provide valuable insights for the referring video segmentation field. We provide the code and updated models at https://github.com/kumuji/sa2va-i
format Preprint
id arxiv_https___arxiv_org_abs_2509_19082
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
Nekrasov, Alexey
Athar, Ali
de Geus, Daan
Hermans, Alexander
Leibe, Bastian
Computer Vision and Pattern Recognition
Sa2VA is a recent model for language-guided dense grounding in images and video that achieves state-of-the-art results on multiple segmentation benchmarks and that has become widely popular. However, we found that Sa2VA does not perform according to its full potential for referring video object segmentation tasks. We identify inconsistencies between training and inference procedures as the key factor holding it back. To mitigate this issue, we propose an improved version of Sa2VA, Sa2VA-i, that rectifies these issues and improves the results. In fact, Sa2VA-i sets a new state of the art for multiple video benchmarks and achieves improvements of up to +11.6 J&F on MeViS, +1.4 on Ref-YT-VOS, +3.3 on Ref-DAVIS and +4.1 on ReVOS using the same Sa2VA checkpoints. With our fixes, the Sa2VA-i-1B model even performs on par with the original Sa2VA-26B model on the MeViS benchmark. We hope that this work will show the importance of seemingly trivial implementation details and that it will provide valuable insights for the referring video segmentation field. We provide the code and updated models at https://github.com/kumuji/sa2va-i
title Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.19082