Timeline and Boundary Guided Diffusion Network for Video Shadow Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Haipeng, Wang, Honqiu, Ye, Tian, Xing, Zhaohu, Ma, Jun, Li, Ping, Wang, Qiong, Zhu, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909293096206336
author Zhou, Haipeng
Wang, Honqiu
Ye, Tian
Xing, Zhaohu
Ma, Jun
Li, Ping
Wang, Qiong
Zhu, Lei
author_facet Zhou, Haipeng
Wang, Honqiu
Ye, Tian
Xing, Zhaohu
Ma, Jun
Li, Ping
Wang, Qiong
Zhu, Lei
contents Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11785
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Timeline and Boundary Guided Diffusion Network for Video Shadow Detection
Zhou, Haipeng
Wang, Honqiu
Ye, Tian
Xing, Zhaohu
Ma, Jun
Li, Ping
Wang, Qiong
Zhu, Lei
Computer Vision and Pattern Recognition
Artificial Intelligence
Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}.
title Timeline and Boundary Guided Diffusion Network for Video Shadow Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2408.11785