Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Linge, Chen, Yingying, Zhu, Bingke, Zhou, Lu, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2024)
von: Wang, Langyu, et al.
Veröffentlicht: (2024)
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
von: Wang, Langyu, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
Semantic Audio-Visual Navigation in Continuous Environments
von: Zeng, Yichen, et al.
Veröffentlicht: (2026)
von: Zeng, Yichen, et al.
Veröffentlicht: (2026)
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
von: Cheng, Hao, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
MathPhys-Guided Coarse-to-Fine Anomaly Synthesis with SQE-Driven Bi-Level Optimization for Anomaly Detection
von: Qian, Long, et al.
Veröffentlicht: (2025)
von: Qian, Long, et al.
Veröffentlicht: (2025)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2024)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2024)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
Semantics-Aware Human Motion Generation from Audio Instructions
von: Wang, Zi-An, et al.
Veröffentlicht: (2025)
von: Wang, Zi-An, et al.
Veröffentlicht: (2025)
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation
von: Zhou, Yupeng, et al.
Veröffentlicht: (2026)
von: Zhou, Yupeng, et al.
Veröffentlicht: (2026)
Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
von: Qian, Long, et al.
Veröffentlicht: (2025)
von: Qian, Long, et al.
Veröffentlicht: (2025)
Segment Beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation
von: Wu, Renjie, et al.
Veröffentlicht: (2023)
von: Wu, Renjie, et al.
Veröffentlicht: (2023)
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
von: Li, Kai, et al.
Veröffentlicht: (2025)
von: Li, Kai, et al.
Veröffentlicht: (2025)
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
von: Bai, Detao, et al.
Veröffentlicht: (2025)
von: Bai, Detao, et al.
Veröffentlicht: (2025)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
von: Liao, Junchao, et al.
Veröffentlicht: (2026)
Leveraging Audio Representations for Vibration-Based Crowd Monitoring in Stadiums
von: Chang, Yen Cheng, et al.
Veröffentlicht: (2025)
von: Chang, Yen Cheng, et al.
Veröffentlicht: (2025)
Dual Mean-Teacher: An Unbiased Semi-Supervised Framework for Audio-Visual Source Localization
von: Guo, Yuxin, et al.
Veröffentlicht: (2024)
von: Guo, Yuxin, et al.
Veröffentlicht: (2024)
AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
WavFlow: Audio Generation in Waveform Space
von: Zhou, Feiyan, et al.
Veröffentlicht: (2026)
von: Zhou, Feiyan, et al.
Veröffentlicht: (2026)
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
von: Chen, Tianle, et al.
Veröffentlicht: (2026)
von: Chen, Tianle, et al.
Veröffentlicht: (2026)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
von: Nakada, Shota, et al.
Veröffentlicht: (2024)
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits
von: Li, Kai, et al.
Veröffentlicht: (2022)
von: Li, Kai, et al.
Veröffentlicht: (2022)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
von: Liu, Zhan, et al.
Veröffentlicht: (2026)
von: Liu, Zhan, et al.
Veröffentlicht: (2026)
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
von: Cheng, Luyao, et al.
Veröffentlicht: (2024)
von: Cheng, Luyao, et al.
Veröffentlicht: (2024)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
von: Xu, Le, et al.
Veröffentlicht: (2025)
von: Xu, Le, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2024) -
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
von: Wang, Langyu, et al.
Veröffentlicht: (2025) -
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
von: Liu, Chen, et al.
Veröffentlicht: (2025) -
Semantic Audio-Visual Navigation in Continuous Environments
von: Zeng, Yichen, et al.
Veröffentlicht: (2026) -
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2024)