MASSeg : 2nd Technical Report for 4th PVUW MOSE Track

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Xuqiang, Zhao, Linnan, Zhao, Jiaxuan, Liu, Fang, Chen, Puhua, Ma, Wenping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908318212030464
author Cao, Xuqiang
Zhao, Linnan
Zhao, Jiaxuan
Liu, Fang
Chen, Puhua
Ma, Wenping
author_facet Cao, Xuqiang
Zhao, Linnan
Zhao, Jiaxuan
Liu, Fang
Chen, Puhua
Ma, Wenping
contents Complex video object segmentation continues to face significant challenges in small object recognition, occlusion handling, and dynamic scene modeling. This report presents our solution, which ranked second in the MOSE track of CVPR 2025 PVUW Challenge. Based on an existing segmentation framework, we propose an improved model named MASSeg for complex video object segmentation, and construct an enhanced dataset, MOSE+, which includes typical scenarios with occlusions, cluttered backgrounds, and small target instances. During training, we incorporate a combination of inter-frame consistent and inconsistent data augmentation strategies to improve robustness and generalization. During inference, we design a mask output scaling strategy to better adapt to varying object sizes and occlusion levels. As a result, MASSeg achieves a J score of 0.8250, F score of 0.9007, and a J&F score of 0.8628 on the MOSE test set.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10254
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
Cao, Xuqiang
Zhao, Linnan
Zhao, Jiaxuan
Liu, Fang
Chen, Puhua
Ma, Wenping
Computer Vision and Pattern Recognition
Artificial Intelligence
Complex video object segmentation continues to face significant challenges in small object recognition, occlusion handling, and dynamic scene modeling. This report presents our solution, which ranked second in the MOSE track of CVPR 2025 PVUW Challenge. Based on an existing segmentation framework, we propose an improved model named MASSeg for complex video object segmentation, and construct an enhanced dataset, MOSE+, which includes typical scenarios with occlusions, cluttered backgrounds, and small target instances. During training, we incorporate a combination of inter-frame consistent and inconsistent data augmentation strategies to improve robustness and generalization. During inference, we design a mask output scaling strategy to better adapt to varying object sizes and occlusion levels. As a result, MASSeg achieves a J score of 0.8250, F score of 0.9007, and a J&F score of 0.8628 on the MOSE test set.
title MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.10254