Reinforced Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Yongbiao, Sun, Xiangcheng, Lv, Guohua, Yu, Deng, Niu, Sijiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Weakly-Supervised Audio-Visual Video Parsing via Segment-wise Pseudo Labeling
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025)
by: Wang, Langyu, et al.
Published: (2025)
CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing
by: Sardari, Faegheh, et al.
Published: (2024)
by: Sardari, Faegheh, et al.
Published: (2024)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
by: Lai, Yung-Hsuan, et al.
Published: (2025)
by: Lai, Yung-Hsuan, et al.
Published: (2025)
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2024)
by: Wang, Langyu, et al.
Published: (2024)
Learning Weakly Supervised Audio-Visual Violence Detection in Hyperbolic Space
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
Cross Pseudo Labeling For Weakly Supervised Video Anomaly Detection
by: Lee, Dayeon, et al.
Published: (2026)
by: Lee, Dayeon, et al.
Published: (2026)
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
by: Zhao, Pengcheng, et al.
Published: (2024)
by: Zhao, Pengcheng, et al.
Published: (2024)
DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing
by: Wang, Xiaoqin, et al.
Published: (2025)
by: Wang, Xiaoqin, et al.
Published: (2025)
Semantic Parsing of Colonoscopy Videos with Multi-Label Temporal Networks
by: Kelner, Ori, et al.
Published: (2023)
by: Kelner, Ori, et al.
Published: (2023)
Leveraging Multi-View Weak Supervision for Occlusion-Aware Multi-Human Parsing
by: Bragagnolo, Laura, et al.
Published: (2025)
by: Bragagnolo, Laura, et al.
Published: (2025)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning
by: Pu, Haojie, et al.
Published: (2026)
by: Pu, Haojie, et al.
Published: (2026)
Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing
by: Song, Jiahe, et al.
Published: (2026)
by: Song, Jiahe, et al.
Published: (2026)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
Learning Event Completeness for Weakly Supervised Video Anomaly Detection
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Weakly-Supervised Semantic Segmentation with Image-Level Labels: from Traditional Models to Foundation Models
by: Chen, Zhaozheng, et al.
Published: (2023)
by: Chen, Zhaozheng, et al.
Published: (2023)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision
by: Gurbuz, A. Said, et al.
Published: (2026)
by: Gurbuz, A. Said, et al.
Published: (2026)
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
by: Wu, Jianzong, et al.
Published: (2025)
by: Wu, Jianzong, et al.
Published: (2025)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026)
by: Shi, Miaojing, et al.
Published: (2026)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
Emerging Trends in Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation with Image-Level Supervision
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
WS-IMUBench: Can Weakly Supervised Methods from Audio, Image, and Video Be Adapted for IMU-based Temporal Action Localization?
by: Li, Pei, et al.
Published: (2026)
by: Li, Pei, et al.
Published: (2026)
Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising
by: Ji, Mingjie, et al.
Published: (2026)
by: Ji, Mingjie, et al.
Published: (2026)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
by: Huang, Yanxiang, et al.
Published: (2026)
by: Huang, Yanxiang, et al.
Published: (2026)
Weakly Supervised Video Scene Graph Generation via Natural Language Supervision
by: Kim, Kibum, et al.
Published: (2025)
by: Kim, Kibum, et al.
Published: (2025)
Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic Alignment
by: Yin, Wenti, et al.
Published: (2025)
by: Yin, Wenti, et al.
Published: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
Enhancing Monocular Height Estimation via Weak Supervision from Imperfect Labels
by: Chen, Sining, et al.
Published: (2025)
by: Chen, Sining, et al.
Published: (2025)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
by: Deng, Yuchen, et al.
Published: (2025)
by: Deng, Yuchen, et al.
Published: (2025)
Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection
by: Leng, Jiaxu, et al.
Published: (2024)
by: Leng, Jiaxu, et al.
Published: (2024)
Similar Items
-
Advancing Weakly-Supervised Audio-Visual Video Parsing via Segment-wise Pseudo Labeling
by: Zhou, Jinxing, et al.
Published: (2024) -
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026) -
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
by: Wang, Langyu, et al.
Published: (2025) -
CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing
by: Sardari, Faegheh, et al.
Published: (2024) -
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
by: Chen, Yaru, et al.
Published: (2025)