Metric Learning with Progressive Self-Distillation for Audio-Visual Embedding Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Donghuo, Ikeda, Kazushi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
by: Zeng, Donghuo, et al.
Published: (2024)
by: Zeng, Donghuo, et al.
Published: (2024)
Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
by: Lin, Junan, et al.
Published: (2025)
by: Lin, Junan, et al.
Published: (2025)
A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
by: Vilaca, Luis, et al.
Published: (2024)
by: Vilaca, Luis, et al.
Published: (2024)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
by: Kang, Minjae, et al.
Published: (2025)
by: Kang, Minjae, et al.
Published: (2025)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
by: Krishnamurthy, Sudha
Published: (2024)
by: Krishnamurthy, Sudha
Published: (2024)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
by: Joo, Seohyun, et al.
Published: (2026)
by: Joo, Seohyun, et al.
Published: (2026)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
Multimodal Transformer Distillation for Audio-Visual Synchronization
by: Chen, Xuanjun, et al.
Published: (2022)
by: Chen, Xuanjun, et al.
Published: (2022)
Audio-Visual Segmentation via Unlabeled Frame Exploitation
by: Liu, Jinxiang, et al.
Published: (2024)
by: Liu, Jinxiang, et al.
Published: (2024)
Object-AVEdit: An Object-level Audio-Visual Editing Model
by: Fu, Youquan, et al.
Published: (2025)
by: Fu, Youquan, et al.
Published: (2025)
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
by: Xie, Liuyue, et al.
Published: (2025)
by: Xie, Liuyue, et al.
Published: (2025)
Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
by: Katamneni, Vinaya Sree, et al.
Published: (2024)
by: Katamneni, Vinaya Sree, et al.
Published: (2024)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
by: Choi, Hahyeon, et al.
Published: (2025)
by: Choi, Hahyeon, et al.
Published: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification
by: Xiao, Zhenyuan, et al.
Published: (2024)
by: Xiao, Zhenyuan, et al.
Published: (2024)
SAV-SE: Scene-aware Audio-Visual Speech Enhancement with Selective State Space Model
by: Qian, Xinyuan, et al.
Published: (2024)
by: Qian, Xinyuan, et al.
Published: (2024)
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
by: Nakada, Shota, et al.
Published: (2024)
by: Nakada, Shota, et al.
Published: (2024)
Enhancing Lie Detection Accuracy: A Comparative Study of Classic ML, CNN, and GCN Models using Audio-Visual Features
by: Abdelwahab, Abdelrahman, et al.
Published: (2024)
by: Abdelwahab, Abdelrahman, et al.
Published: (2024)
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
by: Yu, RunLin, et al.
Published: (2024)
by: Yu, RunLin, et al.
Published: (2024)
Multimodal Sentiment Analysis based on Video and Audio Inputs
by: Fernandez, Antonio, et al.
Published: (2024)
by: Fernandez, Antonio, et al.
Published: (2024)
Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation
by: Liang, Yingshan, et al.
Published: (2025)
by: Liang, Yingshan, et al.
Published: (2025)
TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
by: Choi, Keunwoo, et al.
Published: (2025)
by: Choi, Keunwoo, et al.
Published: (2025)
On the Effect of Data-Augmentation on Local Embedding Properties in the Contrastive Learning of Music Audio Representations
by: McCallum, Matthew C., et al.
Published: (2024)
by: McCallum, Matthew C., et al.
Published: (2024)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
by: Qiang, Chunyu, et al.
Published: (2026)
by: Qiang, Chunyu, et al.
Published: (2026)
Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion
by: Sun, Yu, et al.
Published: (2025)
by: Sun, Yu, et al.
Published: (2025)
Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network
by: Wu, Junyan, et al.
Published: (2025)
by: Wu, Junyan, et al.
Published: (2025)
Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations
by: Yeo, Jeong Hun, et al.
Published: (2025)
by: Yeo, Jeong Hun, et al.
Published: (2025)
3D Audio-Visual Segmentation
by: Sokolov, Artem, et al.
Published: (2024)
by: Sokolov, Artem, et al.
Published: (2024)
MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
CineSRD: Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization
by: Huang, Liangbin, et al.
Published: (2026)
by: Huang, Liangbin, et al.
Published: (2026)
Audio-Vision Contrastive Learning for Phonological Class Recognition
by: Liu, Daiqi, et al.
Published: (2025)
by: Liu, Daiqi, et al.
Published: (2025)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
by: Liu, Zehua, et al.
Published: (2024)
by: Liu, Zehua, et al.
Published: (2024)
Benchmarking Cross-Domain Audio-Visual Deception Detection
by: Guo, Xiaobao, et al.
Published: (2024)
by: Guo, Xiaobao, et al.
Published: (2024)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
by: Xu, Le, et al.
Published: (2025)
by: Xu, Le, et al.
Published: (2025)
Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
by: Astrid, Marcella, et al.
Published: (2024)
by: Astrid, Marcella, et al.
Published: (2024)
Dance Any Beat: Blending Beats with Visuals in Dance Video Generation
by: Wang, Xuanchen, et al.
Published: (2024)
by: Wang, Xuanchen, et al.
Published: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
by: Mo, Shentong, et al.
Published: (2024)
by: Mo, Shentong, et al.
Published: (2024)
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
by: Oorloff, Trevine, et al.
Published: (2024)
by: Oorloff, Trevine, et al.
Published: (2024)
Similar Items
-
Anchor-aware Deep Metric Learning for Audio-visual Retrieval
by: Zeng, Donghuo, et al.
Published: (2024) -
Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval
by: Lin, Junan, et al.
Published: (2025) -
A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
by: Vilaca, Luis, et al.
Published: (2024) -
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
by: Kang, Minjae, et al.
Published: (2025) -
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
by: Krishnamurthy, Sudha
Published: (2024)