A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Guohuan, Hesham, Syed Ariff Syed, Guo, Wenya, Li, Bing, Cheng, Ming-Ming, Sun, Guolei, Liu, Yun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908408956846080
author Xie, Guohuan
Hesham, Syed Ariff Syed
Guo, Wenya
Li, Bing
Cheng, Ming-Ming
Sun, Guolei
Liu, Yun
author_facet Xie, Guohuan
Hesham, Syed Ariff Syed
Guo, Wenya
Li, Bing
Cheng, Ming-Ming
Sun, Guolei
Liu, Yun
contents Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a holistic review of recent advances in VSP, covering a wide array of vision tasks, including Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), as well as Video Tracking and Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We systematically analyze the evolution from traditional hand-crafted features to modern deep learning paradigms -- spanning from fully convolutional networks to the latest transformer-based architectures -- and assess their effectiveness in capturing both local and global temporal contexts. Furthermore, our review critically discusses the technical challenges, ranging from maintaining temporal consistency to handling complex scene dynamics, and offers a comprehensive comparative study of datasets and evaluation metrics that have shaped current benchmarking standards. By distilling the key contributions and shortcomings of state-of-the-art methodologies, this survey highlights emerging trends and prospective research directions that promise to further elevate the robustness and adaptability of VSP in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13552
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
Xie, Guohuan
Hesham, Syed Ariff Syed
Guo, Wenya
Li, Bing
Cheng, Ming-Ming
Sun, Guolei
Liu, Yun
Computer Vision and Pattern Recognition
Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a holistic review of recent advances in VSP, covering a wide array of vision tasks, including Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), as well as Video Tracking and Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We systematically analyze the evolution from traditional hand-crafted features to modern deep learning paradigms -- spanning from fully convolutional networks to the latest transformer-based architectures -- and assess their effectiveness in capturing both local and global temporal contexts. Furthermore, our review critically discusses the technical challenges, ranging from maintaining temporal consistency to handling complex scene dynamics, and offers a comprehensive comparative study of datasets and evaluation metrics that have shaped current benchmarking standards. By distilling the key contributions and shortcomings of state-of-the-art methodologies, this survey highlights emerging trends and prospective research directions that promise to further elevate the robustness and adaptability of VSP in real-world applications.
title A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13552