CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yuanhong, Shimada, Kazuki, Simon, Christian, Ikemiya, Yukara, Shibuya, Takashi, Mitsufuji, Yuki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
StereoSync: Spatially-Aware Stereo Audio Generation from Video
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025)
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
von: Hayakawa, Akio, et al.
Veröffentlicht: (2024)
von: Hayakawa, Akio, et al.
Veröffentlicht: (2024)
Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
von: Yang, Shiqi, et al.
Veröffentlicht: (2024)
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
von: Cheng, Ho Kei, et al.
Veröffentlicht: (2024)
von: Cheng, Ho Kei, et al.
Veröffentlicht: (2024)
A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
Zero- and Few-shot Sound Event Localization and Detection
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
Variable Bitrate Residual Vector Quantization for Audio Coding
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
von: Chae, Yunkee, et al.
Veröffentlicht: (2024)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
von: Takahashi, Akira, et al.
Veröffentlicht: (2025)
von: Takahashi, Akira, et al.
Veröffentlicht: (2025)
Deep Learning for Personalized Binaural Audio Reproduction
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
von: Bai, Detao, et al.
Veröffentlicht: (2025)
von: Bai, Detao, et al.
Veröffentlicht: (2025)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
Binamix -- A Python Library for Generating Binaural Audio Datasets
von: Barry, Dan, et al.
Veröffentlicht: (2025)
von: Barry, Dan, et al.
Veröffentlicht: (2025)
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
Lightweight Implicit Neural Network for Binaural Audio Synthesis
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
von: Lu, Xikun, et al.
Veröffentlicht: (2025)
Sequential Contrastive Audio-Visual Learning
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders
von: Shi, Hao, et al.
Veröffentlicht: (2023)
von: Shi, Hao, et al.
Veröffentlicht: (2023)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio
von: Chen, Gongyu, et al.
Veröffentlicht: (2024)
von: Chen, Gongyu, et al.
Veröffentlicht: (2024)
DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
Audio-Vision Contrastive Learning for Phonological Class Recognition
von: Liu, Daiqi, et al.
Veröffentlicht: (2025)
von: Liu, Daiqi, et al.
Veröffentlicht: (2025)
OmniAudio: Generating Spatial Audio from 360-Degree Video
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
von: Panah, Davoud Shariat, et al.
Veröffentlicht: (2025)
von: Panah, Davoud Shariat, et al.
Veröffentlicht: (2025)
MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment
von: Zhou, Hao, et al.
Veröffentlicht: (2025)
von: Zhou, Hao, et al.
Veröffentlicht: (2025)
SilentCipher: Deep Audio Watermarking
von: Singh, Mayank Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Mayank Kumar, et al.
Veröffentlicht: (2024)
TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
Learning to Highlight Audio by Watching Movies
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
UniSync: A Unified Framework for Audio-Visual Synchronization
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
GRAFX: An Open-Source Library for Audio Processing Graphs in PyTorch
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
von: Lee, Sungho, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
von: Shimada, Kazuki, et al.
Veröffentlicht: (2024) -
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024) -
StereoSync: Spatially-Aware Stereo Audio Generation from Video
von: Marinoni, Christian, et al.
Veröffentlicht: (2025) -
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
von: Hayakawa, Akio, et al.
Veröffentlicht: (2025) -
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
von: Hayakawa, Akio, et al.
Veröffentlicht: (2024)