A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Guohuan, Hesham, Syed Ariff Syed, Guo, Wenya, Li, Bing, Cheng, Ming-Ming, Sun, Guolei, Liu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating SAM2 for Video Semantic Segmentation
by: Ariff, Syed Hesham Syed, et al.
Published: (2025)
by: Ariff, Syed Hesham Syed, et al.
Published: (2025)
RGB-D Indiscernible Object Counting in Underwater Scenes
by: Sun, Guolei, et al.
Published: (2023)
by: Sun, Guolei, et al.
Published: (2023)
Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation
by: Xie, Guohuan, et al.
Published: (2026)
by: Xie, Guohuan, et al.
Published: (2026)
Traffic Scene Parsing through the TSP6K Dataset
by: Jiang, Peng-Tao, et al.
Published: (2023)
by: Jiang, Peng-Tao, et al.
Published: (2023)
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
by: An, Zhaochong, et al.
Published: (2024)
by: An, Zhaochong, et al.
Published: (2024)
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
by: Zhou, Yuli, et al.
Published: (2024)
by: Zhou, Yuli, et al.
Published: (2024)
A Comprehensive Survey for Real-World Industrial Defect Detection: Challenges, Approaches, and Prospects
by: Cheng, Yuqi, et al.
Published: (2025)
by: Cheng, Yuqi, et al.
Published: (2025)
Advances in Deep Concealed Scene Understanding
by: Fan, Deng-Ping, et al.
Published: (2023)
by: Fan, Deng-Ping, et al.
Published: (2023)
PIG: Prompt Images Guidance for Night-Time Scene Parsing
by: Xie, Zhifeng, et al.
Published: (2024)
by: Xie, Zhifeng, et al.
Published: (2024)
A Survey on Long Video Generation: Challenges, Methods, and Prospects
by: Li, Chengxuan, et al.
Published: (2024)
by: Li, Chengxuan, et al.
Published: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
Hypergraph-Enhanced Training-Free and Language-Free Few-Shot Anomaly Detection
by: Xie, Guohuan, et al.
Published: (2026)
by: Xie, Guohuan, et al.
Published: (2026)
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing
by: Fan, Jiahe, et al.
Published: (2026)
by: Fan, Jiahe, et al.
Published: (2026)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
by: Sun, Guolei, et al.
Published: (2022)
by: Sun, Guolei, et al.
Published: (2022)
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
by: Xu, Pengxin, et al.
Published: (2026)
by: Xu, Pengxin, et al.
Published: (2026)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025)
by: Wang, Langyu, et al.
Published: (2025)
MedVSR: Medical Video Super-Resolution with Cross State-Space Propagation
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Fine-Grained Zero-Shot Learning: Advances, Challenges, and Prospects
by: Guo, Jingcai, et al.
Published: (2024)
by: Guo, Jingcai, et al.
Published: (2024)
Advancing Weakly-Supervised Audio-Visual Video Parsing via Segment-wise Pseudo Labeling
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
AceVFI: A Comprehensive Survey of Advances in Video Frame Interpolation
by: Kye, Dahyeon, et al.
Published: (2025)
by: Kye, Dahyeon, et al.
Published: (2025)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware Information Decoupling and Advanced Heterogeneous Feature Fusion
by: Huang, Jianxin, et al.
Published: (2024)
by: Huang, Jianxin, et al.
Published: (2024)
HAPNet: Toward Superior RGB-Thermal Scene Parsing via Hybrid, Asymmetric, and Progressive Heterogeneous Feature Fusion
by: Li, Jiahang, et al.
Published: (2024)
by: Li, Jiahang, et al.
Published: (2024)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
by: Li, Jiahang, et al.
Published: (2023)
by: Li, Jiahang, et al.
Published: (2023)
CoCo4D: Comprehensive and Complex 4D Scene Generation
by: Zhou, Junwei, et al.
Published: (2025)
by: Zhou, Junwei, et al.
Published: (2025)
DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes
by: Liu, Jinxiu, et al.
Published: (2024)
by: Liu, Jinxiu, et al.
Published: (2024)
Improvement of Human-Object Interaction Action Recognition Using Scene Information and Multi-Task Learning Approach
by: Shehata, Hesham M., et al.
Published: (2025)
by: Shehata, Hesham M., et al.
Published: (2025)
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
by: Tariq, Omer, et al.
Published: (2026)
by: Tariq, Omer, et al.
Published: (2026)
DINO-Mix: Distilling Foundational Knowledge with Cross-Domain CutMix for Semi-supervised Class-imbalanced Medical Image Segmentation
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
ViTs for Action Classification in Videos: An Approach to Risky Tackle Detection in American Football Practice Videos
by: Zaidi, Syed Ahsan Masud, et al.
Published: (2026)
by: Zaidi, Syed Ahsan Masud, et al.
Published: (2026)
A Survey on Ordinal Regression: Applications, Advances and Prospects
by: Wang, Jinhong, et al.
Published: (2025)
by: Wang, Jinhong, et al.
Published: (2025)
Visibility-Uncertainty-guided 3D Gaussian Inpainting via Scene Conceptional Learning
by: Cui, Mingxuan, et al.
Published: (2025)
by: Cui, Mingxuan, et al.
Published: (2025)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
by: Bao, Jingzhi, et al.
Published: (2024)
by: Bao, Jingzhi, et al.
Published: (2024)
RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
by: Sun, Zhichao, et al.
Published: (2025)
by: Sun, Zhichao, et al.
Published: (2025)
Video Unsupervised Domain Adaptation with Deep Learning: A Comprehensive Survey
by: Xu, Yuecong, et al.
Published: (2022)
by: Xu, Yuecong, et al.
Published: (2022)
Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
by: Gao, Yongbiao, et al.
Published: (2024)
by: Gao, Yongbiao, et al.
Published: (2024)
A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights
by: Lei, Wentao, et al.
Published: (2024)
by: Lei, Wentao, et al.
Published: (2024)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
by: Zhang, Qintong, et al.
Published: (2024)
by: Zhang, Qintong, et al.
Published: (2024)
Advancing Textual Prompt Learning with Anchored Attributes
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
Similar Items
-
Evaluating SAM2 for Video Semantic Segmentation
by: Ariff, Syed Hesham Syed, et al.
Published: (2025) -
RGB-D Indiscernible Object Counting in Underwater Scenes
by: Sun, Guolei, et al.
Published: (2023) -
Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation
by: Xie, Guohuan, et al.
Published: (2026) -
Traffic Scene Parsing through the TSP6K Dataset
by: Jiang, Peng-Tao, et al.
Published: (2023) -
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
by: An, Zhaochong, et al.
Published: (2024)