Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915499580850176 |
|---|---|
| author | Chen, Yaru Guo, Ruohao Gao, Liting Xiang, Yang Luo, Qingyu Li, Zhenbo Wang, Wenwu |
| author_facet | Chen, Yaru Guo, Ruohao Gao, Liting Xiang, Yang Luo, Qingyu Li, Zhenbo Wang, Wenwu |
| contents | Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative learning, but neglected stable segment-level supervision and class-aware cross-modal alignment. To address this, we propose two strategies: (1) an exponential moving average (EMA)-guided pseudo supervision framework that generates reliable segment-level masks via adaptive thresholds or top-k selection, offering stable temporal guidance beyond video-level labels; and (2) a class-aware cross-modal agreement (CMA) loss that aligns audio and visual embeddings at reliable segment-class pairs, ensuring consistency across modalities while preserving temporal structure. Evaluations on LLP and UnAV-100 datasets shows that our method achieves state-of-the-art (SOTA) performance across multiple metrics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14097 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing Chen, Yaru Guo, Ruohao Gao, Liting Xiang, Yang Luo, Qingyu Li, Zhenbo Wang, Wenwu Computer Vision and Pattern Recognition Multimedia Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative learning, but neglected stable segment-level supervision and class-aware cross-modal alignment. To address this, we propose two strategies: (1) an exponential moving average (EMA)-guided pseudo supervision framework that generates reliable segment-level masks via adaptive thresholds or top-k selection, offering stable temporal guidance beyond video-level labels; and (2) a class-aware cross-modal agreement (CMA) loss that aligns audio and visual embeddings at reliable segment-class pairs, ensuring consistency across modalities while preserving temporal structure. Evaluations on LLP and UnAV-100 datasets shows that our method achieves state-of-the-art (SOTA) performance across multiple metrics. |
| title | Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing |
| topic | Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2509.14097 |