Woosh: A Sound Effects Foundation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hadjeres, Gaëtan, Ferras, Marc, Koutini, Khaled, Weck, Benno, Bittar, Alexandre, Hummel, Thomas, Lahrichi, Zineb, Missoum, Hakim, Serrà, Joan, Mitsufuji, Yuki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026)
QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2025)
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2025)
Automatic Music Sample Identification with Multi-Track Contrastive Learning
von: Riou, Alain, et al.
Veröffentlicht: (2025)
von: Riou, Alain, et al.
Veröffentlicht: (2025)
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2023)
von: Riou, Alain, et al.
Veröffentlicht: (2023)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
von: Weck, Benno, et al.
Veröffentlicht: (2023)
von: Weck, Benno, et al.
Veröffentlicht: (2023)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2025)
The language of sound search: Examining User Queries in Audio Search Engines
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
von: Morreale, Fabio, et al.
Veröffentlicht: (2025)
von: Morreale, Fabio, et al.
Veröffentlicht: (2025)
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
von: Serrà, Joan, et al.
Veröffentlicht: (2025)
von: Serrà, Joan, et al.
Veröffentlicht: (2025)
Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning
von: Riou, Alain, et al.
Veröffentlicht: (2024)
von: Riou, Alain, et al.
Veröffentlicht: (2024)
Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
HumMusQA: A Human-written Music Understanding QA Benchmark Dataset
von: Weck, Benno, et al.
Veröffentlicht: (2026)
von: Weck, Benno, et al.
Veröffentlicht: (2026)
Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures
von: Riou, Alain, et al.
Veröffentlicht: (2024)
von: Riou, Alain, et al.
Veröffentlicht: (2024)
Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
von: Riou, Alain, et al.
Veröffentlicht: (2024)
von: Riou, Alain, et al.
Veröffentlicht: (2024)
'Studies for': A Human-AI Co-Creative Sound Artwork Using a Real-time Multi-channel Sound Generation Model
von: Nagashima, Chihiro, et al.
Veröffentlicht: (2025)
von: Nagashima, Chihiro, et al.
Veröffentlicht: (2025)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
Automatic Music Mixing using a Generative Model of Effect Embeddings
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
von: Takahashi, Akira, et al.
Veröffentlicht: (2025)
von: Takahashi, Akira, et al.
Veröffentlicht: (2025)
A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
Zero- and Few-shot Sound Event Localization and Detection
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
Do Foundational Audio Encoders Understand Music Structure?
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
Data-Efficient Low-Complexity Acoustic Scene Classification in the DCASE 2024 Challenge
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
von: Tsoi, Tristan, et al.
Veröffentlicht: (2025)
von: Tsoi, Tristan, et al.
Veröffentlicht: (2025)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2024)
von: Saito, Koichi, et al.
Veröffentlicht: (2024)
A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
von: Özer, Yigitcan, et al.
Veröffentlicht: (2025)
von: Özer, Yigitcan, et al.
Veröffentlicht: (2025)
PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2025)
von: Riou, Alain, et al.
Veröffentlicht: (2025)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
von: Weck, Benno, et al.
Veröffentlicht: (2024)
von: Weck, Benno, et al.
Veröffentlicht: (2024)
Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits
von: Ishii, Masato, et al.
Veröffentlicht: (2025)
von: Ishii, Masato, et al.
Veröffentlicht: (2025)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
von: Choi, Woosung, et al.
Veröffentlicht: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
Diffusion-based Signal Refiner for Speech Enhancement and Separation
von: Hirano, Masato, et al.
Veröffentlicht: (2023)
von: Hirano, Masato, et al.
Veröffentlicht: (2023)
The Whole Is Greater than the Sum of Its Parts: Improving Music Source Separation by Bridging Network
von: Sawata, Ryosuke, et al.
Veröffentlicht: (2023)
von: Sawata, Ryosuke, et al.
Veröffentlicht: (2023)
ITO-Master: Inference-Time Optimization for Audio Effects Modeling of Music Mastering Processors
von: Koo, Junghyun, et al.
Veröffentlicht: (2025)
von: Koo, Junghyun, et al.
Veröffentlicht: (2025)
SilentCipher: Deep Audio Watermarking
von: Singh, Mayank Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Mayank Kumar, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2026) -
QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
von: Lahrichi, Zineb, et al.
Veröffentlicht: (2025) -
Automatic Music Sample Identification with Multi-Track Contrastive Learning
von: Riou, Alain, et al.
Veröffentlicht: (2025) -
PESTO: Pitch Estimation with Self-supervised Transposition-equivariant Objective
von: Riou, Alain, et al.
Veröffentlicht: (2023) -
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
von: Weck, Benno, et al.
Veröffentlicht: (2023)