Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jisheng, Dang, Xudong, Wu, Bimei, Wang, Ning, Lv, Jiayu, Chen, Zhao, Jingwen, liu, Yichu, Liu, Jizhao, Li, Juncheng, Wang, Teng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915363955933184
author Jisheng, Dang
Xudong, Wu
Bimei, Wang
Ning, Lv
Jiayu, Chen
Zhao, Jingwen
liu, Yichu
Liu, Jizhao
Li, Juncheng
Wang, Teng
author_facet Jisheng, Dang
Xudong, Wu
Bimei, Wang
Ning, Lv
Jiayu, Chen
Zhao, Jingwen
liu, Yichu
Liu, Jizhao
Li, Juncheng
Wang, Teng
contents Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semantics, thereby degrading segmentation accuracy. To systematically mitigate this issue, we propose DeSa2VA, a decoupling-enhanced prompting scheme integrating text pre-training and a linear decoupling module to address the information processing limitations inherent in SAM-2. Specifically, first, we devise a pre-training paradigm that converts textual ground-truth labels into point-level prompts while generating corresponding text masks. These masks are refined through a hybrid loss function to strengthen the model's semantic grounding capabilities. Next, we employ linear projection to disentangle hidden states that generated by a large language model into distinct textual and visual feature subspaces. Finally, a dynamic mask fusion strategy synergistically combines these decoupled features through triple supervision from predicted text/visual masks and ground-truth annotations. Extensive experiments demonstrate state-of-the-art performance across diverse tasks, including image segmentation, image question answering, video segmentation, and video question answering. Our codes are available at https://github.com/longmalongma/DeSa2VA.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22880
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
Jisheng, Dang
Xudong, Wu
Bimei, Wang
Ning, Lv
Jiayu, Chen
Zhao, Jingwen
liu, Yichu
Liu, Jizhao
Li, Juncheng
Wang, Teng
Computer Vision and Pattern Recognition
Artificial Intelligence
Existing video segmenter and grounder approaches, exemplified by Sa2VA, directly fuse features within segmentation models. This often results in an undesirable entanglement of dynamic visual information and static semantics, thereby degrading segmentation accuracy. To systematically mitigate this issue, we propose DeSa2VA, a decoupling-enhanced prompting scheme integrating text pre-training and a linear decoupling module to address the information processing limitations inherent in SAM-2. Specifically, first, we devise a pre-training paradigm that converts textual ground-truth labels into point-level prompts while generating corresponding text masks. These masks are refined through a hybrid loss function to strengthen the model's semantic grounding capabilities. Next, we employ linear projection to disentangle hidden states that generated by a large language model into distinct textual and visual feature subspaces. Finally, a dynamic mask fusion strategy synergistically combines these decoupled features through triple supervision from predicted text/visual masks and ground-truth annotations. Extensive experiments demonstrate state-of-the-art performance across diverse tasks, including image segmentation, image question answering, video segmentation, and video question answering. Our codes are available at https://github.com/longmalongma/DeSa2VA.
title Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.22880