Latent Watermarking of Audio Generative Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Roman, Robin San, Fernandez, Pierre, Deleforge, Antoine, Adi, Yossi, Serizel, Romain |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
Angular Distance Distribution Loss for Audio Classification
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023)
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)
von: Aziz, Shiran, et al.
Veröffentlicht: (2024)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2024)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2024)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
Energy Consumption Trends in Sound Event Detection Systems
von: Douwes, Constance, et al.
Veröffentlicht: (2024)
von: Douwes, Constance, et al.
Veröffentlicht: (2024)
A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
Domain-Invariant Representation Learning of Bird Sounds
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
von: Passoni, Riccardo, et al.
Veröffentlicht: (2025)
von: Passoni, Riccardo, et al.
Veröffentlicht: (2025)
A decade of DCASE: Achievements, practices, evaluations and future challenges
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024)
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024)
Self-Supervised Learning for Few-Shot Bird Sound Classification
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
WAKE: Watermarking Audio with Key Enrichment
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2025)
Fully Reversing the Shoebox Image Source Method: From Impulse Responses to Room Parameters
von: Sprunck, Tom, et al.
Veröffentlicht: (2024)
von: Sprunck, Tom, et al.
Veröffentlicht: (2024)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
Salmon: A Suite for Acoustic Language Model Evaluation
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
Low-Resource Audio Codec (LRAC): 2025 Challenge Description
von: Wojcicki, Kamil, et al.
Veröffentlicht: (2025)
von: Wojcicki, Kamil, et al.
Veröffentlicht: (2025)
Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024)
von: Roth, Amit, et al.
Veröffentlicht: (2024)
CAFA: a Controllable Automatic Foley Artist
von: Benita, Roi, et al.
Veröffentlicht: (2025)
von: Benita, Roi, et al.
Veröffentlicht: (2025)
Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2022)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
The impact of non-target events in synthetic soundscapes for sound event detection
von: Ronchini, Francesca, et al.
Veröffentlicht: (2021)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2021)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
von: Iatariene, Taous, et al.
Veröffentlicht: (2025)
Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
von: Sadeghi, Mostafa, et al.
Veröffentlicht: (2025)
von: Sadeghi, Mostafa, et al.
Veröffentlicht: (2025)
DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
WavMark: Watermarking for Audio Generation
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2025)
von: Tal, Or, et al.
Veröffentlicht: (2025)
WHISTRESS: Enriching Transcriptions with Sentence Stress Detection
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
von: Rouard, Simon, et al.
Veröffentlicht: (2025) -
Angular Distance Distribution Loss for Audio Classification
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024) -
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
von: Ronchini, Francesca, et al.
Veröffentlicht: (2023) -
Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
von: Aziz, Shiran, et al.
Veröffentlicht: (2024) -
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)