Latent CLAP Loss for Better Foley Sound Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Karchkhadze, Tornike, Kavaki, Hassan Salami, Izadi, Mohammad Rasool, Irvin, Bryce, Kegler, Mikolaj, Hertz, Ari, Zhang, Shuo, Stamenovic, Marko |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
CATSE: A Context-Aware Framework for Causal Target Sound Extraction
by: Baligar, Shrishail, et al.
Published: (2024)
by: Baligar, Shrishail, et al.
Published: (2024)
Improving Music Source Separation with Diffusion and Consistency Refinement
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables
by: Ohlenbusch, Mattes, et al.
Published: (2025)
by: Ohlenbusch, Mattes, et al.
Published: (2025)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
FSD50K-Solo: Automated Curation of Single-Source Sound Events
by: Yang, Ningyuan, et al.
Published: (2026)
by: Yang, Ningyuan, et al.
Published: (2026)
StereoFoley: Object-Aware Stereo Audio Generation from Video
by: Karchkhadze, Tornike, et al.
Published: (2025)
by: Karchkhadze, Tornike, et al.
Published: (2025)
Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's Treatise
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
"It is okay to be uncommon": Quantizing Sound Event Detection Networks on Hardware Accelerators with Uncommon Sub-Byte Support
by: Wu, Yushu, et al.
Published: (2024)
by: Wu, Yushu, et al.
Published: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
by: Colombo, Marco Furio, et al.
Published: (2024)
by: Colombo, Marco Furio, et al.
Published: (2024)
HiSSNet: Sound Event Detection and Speaker Identification via Hierarchical Prototypical Networks for Low-Resource Headphones
by: Shashaank, N, et al.
Published: (2023)
by: Shashaank, N, et al.
Published: (2023)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
CAFA: a Controllable Automatic Foley Artist
by: Benita, Roi, et al.
Published: (2025)
by: Benita, Roi, et al.
Published: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
by: Zhang, Yaoyun, et al.
Published: (2024)
by: Zhang, Yaoyun, et al.
Published: (2024)
Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
by: Lee, Junwon, et al.
Published: (2024)
by: Lee, Junwon, et al.
Published: (2024)
Masked Audio Modeling with CLAP and Multi-Objective Learning
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
by: Li, Xiquan, et al.
Published: (2024)
by: Li, Xiquan, et al.
Published: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
by: Kim, Jaeyeon, et al.
Published: (2024)
by: Kim, Jaeyeon, et al.
Published: (2024)
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
CLAP-Based Automatic Word Naming Recognition in Post-Stroke Aphasia
by: Kaloga, Yacouba, et al.
Published: (2026)
by: Kaloga, Yacouba, et al.
Published: (2026)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
by: Takano, Taisei, et al.
Published: (2025)
by: Takano, Taisei, et al.
Published: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
by: Niu, Zhikang, et al.
Published: (2025)
by: Niu, Zhikang, et al.
Published: (2025)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
by: Chu, Annie, et al.
Published: (2024)
by: Chu, Annie, et al.
Published: (2024)
Audio Palette: A Diffusion Transformer with Multi-Signal Conditioning for Controllable Foley Synthesis
by: Wang, Junnuo
Published: (2025)
by: Wang, Junnuo
Published: (2025)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
by: Jing, Xin, et al.
Published: (2024)
by: Jing, Xin, et al.
Published: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
by: Jing, Xin, et al.
Published: (2026)
by: Jing, Xin, et al.
Published: (2026)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
by: Chen, Wenxi, et al.
Published: (2024)
by: Chen, Wenxi, et al.
Published: (2024)
AnoPatch: Towards Better Consistency in Machine Anomalous Sound Detection
by: Jiang, Anbai, et al.
Published: (2024)
by: Jiang, Anbai, et al.
Published: (2024)
The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels
by: Tsutsumi, Ayuto, et al.
Published: (2026)
by: Tsutsumi, Ayuto, et al.
Published: (2026)
FoleyBench: A Benchmark For Video-to-Audio Models
by: Dixit, Satvik, et al.
Published: (2025)
by: Dixit, Satvik, et al.
Published: (2025)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
by: Sun, Haoqin, et al.
Published: (2025)
by: Sun, Haoqin, et al.
Published: (2025)
SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
by: Niu, Xinlei, et al.
Published: (2024)
by: Niu, Xinlei, et al.
Published: (2024)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025)
by: Shan, Sizhe, et al.
Published: (2025)
CLAP-S: Support Set Based Adaptation for Downstream Fiber-optic Acoustic Recognition
by: Sun, Jingchen, et al.
Published: (2025)
by: Sun, Jingchen, et al.
Published: (2025)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
by: Niizumi, Daisuke, et al.
Published: (2024)
by: Niizumi, Daisuke, et al.
Published: (2024)
Zero-Shot Crate Digging: DJ Tool Retrieval Using Speech Activity, Music Structure And CLAP Embeddings
by: Orife, Iroro
Published: (2024)
by: Orife, Iroro
Published: (2024)
SPO-CLAPScore: Enhancing CLAP-based alignment prediction system with Standardize Preference Optimization, for the first XACLE Challenge
by: Takano, Taisei, et al.
Published: (2026)
by: Takano, Taisei, et al.
Published: (2026)
Low-complexity Attention-based Unsupervised Anomalous Sound Detection exploiting Separable Convolutions and Angular Loss
by: Neri, Michael, et al.
Published: (2024)
by: Neri, Michael, et al.
Published: (2024)
M2D-CLAP: Exploring General-purpose Audio-Language Representations Beyond CLAP
by: Niizumi, Daisuke, et al.
Published: (2025)
by: Niizumi, Daisuke, et al.
Published: (2025)
Similar Items
-
Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion Models
by: Karchkhadze, Tornike, et al.
Published: (2024) -
CATSE: A Context-Aware Framework for Causal Target Sound Extraction
by: Baligar, Shrishail, et al.
Published: (2024) -
Improving Music Source Separation with Diffusion and Consistency Refinement
by: Karchkhadze, Tornike, et al.
Published: (2024) -
PAS-SE: Personalized Auxiliary-Sensor Speech Enhancement for Voice Pickup in Hearables
by: Ohlenbusch, Mattes, et al.
Published: (2025) -
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
by: Karchkhadze, Tornike, et al.
Published: (2024)