Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yongbiao, Sun, Xiangcheng, Lv, Guohua, Yu, Deng, Niu, Sijiu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913627669266432
author Gao, Yongbiao
Sun, Xiangcheng
Lv, Guohua
Yu, Deng
Niu, Sijiu
author_facet Gao, Yongbiao
Sun, Xiangcheng
Lv, Guohua
Yu, Deng
Niu, Sijiu
contents Audio-visual video parsing (AVVP) aims to recognize audio and visual event labels with precise temporal boundaries, which is quite challenging since audio or visual modality might include only one event label with only the overall video labels available. Existing label denoising models often treat the denoising process as a separate preprocessing step, leading to a disconnect between label denoising and AVVP tasks. To bridge this gap, we present a novel joint reinforcement learning-based label denoising approach (RLLD). This approach enables simultaneous training of both label denoising and video parsing models through a joint optimization strategy. We introduce a novel AVVP-validation and soft inter-reward feedback mechanism that directly guides the learning of label denoising policy. Extensive experiments on AVVP tasks demonstrate the superior performance of our proposed method compared to label denoising techniques. Furthermore, by incorporating our label denoising method into other AVVP models, we find that it can further enhance parsing results.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19563
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
Gao, Yongbiao
Sun, Xiangcheng
Lv, Guohua
Yu, Deng
Niu, Sijiu
Computer Vision and Pattern Recognition
Audio-visual video parsing (AVVP) aims to recognize audio and visual event labels with precise temporal boundaries, which is quite challenging since audio or visual modality might include only one event label with only the overall video labels available. Existing label denoising models often treat the denoising process as a separate preprocessing step, leading to a disconnect between label denoising and AVVP tasks. To bridge this gap, we present a novel joint reinforcement learning-based label denoising approach (RLLD). This approach enables simultaneous training of both label denoising and video parsing models through a joint optimization strategy. We introduce a novel AVVP-validation and soft inter-reward feedback mechanism that directly guides the learning of label denoising policy. Extensive experiments on AVVP tasks demonstrate the superior performance of our proposed method compared to label denoising techniques. Furthermore, by incorporating our label denoising method into other AVVP models, we find that it can further enhance parsing results.
title Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19563