Combining audio control and style transfer using latent diffusion
Fuente:
arXiv
Salvato in:
| Autori principali: | Demerlé, Nils, Esling, Philippe, Doras, Guillaume, Genova, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Unsupervised Composable Representations for Audio
di: Bindi, Giovanni, et al.
Pubblicazione: (2024)
di: Bindi, Giovanni, et al.
Pubblicazione: (2024)
Long-form music generation with latent diffusion
di: Evans, Zach, et al.
Pubblicazione: (2024)
di: Evans, Zach, et al.
Pubblicazione: (2024)
Sentiment analysis in non-fixed length audios using a Fully Convolutional Neural Network
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
di: Messina, Francisco, et al.
Pubblicazione: (2025)
di: Messina, Francisco, et al.
Pubblicazione: (2025)
Transformation of audio embeddings into interpretable, concept-based representations
di: Zhang, Alice, et al.
Pubblicazione: (2025)
di: Zhang, Alice, et al.
Pubblicazione: (2025)
Towards generalizing deep-audio fake detection networks
di: Gasenzer, Konstantin, et al.
Pubblicazione: (2023)
di: Gasenzer, Konstantin, et al.
Pubblicazione: (2023)
Testing chatbots on the creation of encoders for audio conditioned image generation
di: León, Jorge E., et al.
Pubblicazione: (2025)
di: León, Jorge E., et al.
Pubblicazione: (2025)
Unsupervised outlier detection to improve bird audio dataset labels
di: Collins, Bruce
Pubblicazione: (2025)
di: Collins, Bruce
Pubblicazione: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
di: Jezidžić, Marin, et al.
Pubblicazione: (2024)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
di: Pepino, Leonardo, et al.
Pubblicazione: (2023)
di: Pepino, Leonardo, et al.
Pubblicazione: (2023)
Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
di: Hummel, Hilde I., et al.
Pubblicazione: (2026)
di: Hummel, Hilde I., et al.
Pubblicazione: (2026)
Versatile audio-visual learning for emotion recognition
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
di: Goncalves, Lucas, et al.
Pubblicazione: (2023)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
di: Siuzdak, Hubert
Pubblicazione: (2023)
di: Siuzdak, Hubert
Pubblicazione: (2023)
Real-time implementation of vibrato transfer as an audio effect
di: Hyrkas, Jeremy
Pubblicazione: (2025)
di: Hyrkas, Jeremy
Pubblicazione: (2025)
Approaching an unknown communication system by latent space exploration and causal inference
di: Beguš, Gašper, et al.
Pubblicazione: (2023)
di: Beguš, Gašper, et al.
Pubblicazione: (2023)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
di: Lin, Tsung-En, et al.
Pubblicazione: (2025)
di: Lin, Tsung-En, et al.
Pubblicazione: (2025)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
di: Dang, Shaoxiang, et al.
Pubblicazione: (2024)
di: Dang, Shaoxiang, et al.
Pubblicazione: (2024)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
di: Robinson, David, et al.
Pubblicazione: (2024)
di: Robinson, David, et al.
Pubblicazione: (2024)
Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
di: Lluís, Francesc, et al.
Pubblicazione: (2024)
di: Lluís, Francesc, et al.
Pubblicazione: (2024)
Generalization in birdsong classification: impact of transfer learning methods and dataset characteristics
di: Ghani, Burooj, et al.
Pubblicazione: (2024)
di: Ghani, Burooj, et al.
Pubblicazione: (2024)
FlowDec: A flow-based full-band general audio codec with high perceptual quality
di: Welker, Simon, et al.
Pubblicazione: (2025)
di: Welker, Simon, et al.
Pubblicazione: (2025)
Recomposer: Event-roll-guided generative audio editing
di: Ellis, Daniel P. W., et al.
Pubblicazione: (2025)
di: Ellis, Daniel P. W., et al.
Pubblicazione: (2025)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
Gender-ambiguous voice generation through feminine speaking style transfer in male voices
di: Koutsogiannaki, Maria, et al.
Pubblicazione: (2024)
di: Koutsogiannaki, Maria, et al.
Pubblicazione: (2024)
Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
di: Mohammadi, Amirmohammad, et al.
Pubblicazione: (2024)
di: Mohammadi, Amirmohammad, et al.
Pubblicazione: (2024)
Exploring bat song syllable representations in self-supervised audio encoders
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
Scaling up masked audio encoder learning for general audio classification
di: Dinkel, Heinrich, et al.
Pubblicazione: (2024)
di: Dinkel, Heinrich, et al.
Pubblicazione: (2024)
DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution
di: Pizarro, Matías, et al.
Pubblicazione: (2023)
di: Pizarro, Matías, et al.
Pubblicazione: (2023)
Speaker anonymization using neural audio codec language models
di: Panariello, Michele, et al.
Pubblicazione: (2023)
di: Panariello, Michele, et al.
Pubblicazione: (2023)
Discriminant audio properties in deep learning based respiratory insufficiency detection in Brazilian Portuguese
di: Gauy, Marcelo Matheus, et al.
Pubblicazione: (2024)
di: Gauy, Marcelo Matheus, et al.
Pubblicazione: (2024)
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
di: Serrà, Joan, et al.
Pubblicazione: (2025)
di: Serrà, Joan, et al.
Pubblicazione: (2025)
ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
di: Liu, Yunyi, et al.
Pubblicazione: (2024)
di: Liu, Yunyi, et al.
Pubblicazione: (2024)
Acoustic and Machine Learning Methods for Speech-Based Suicide Risk Assessment: A Systematic Review
di: Marie, Ambre, et al.
Pubblicazione: (2025)
di: Marie, Ambre, et al.
Pubblicazione: (2025)
Episodic fine-tuning prototypical networks for optimization-based few-shot learning: Application to audio classification
di: Zhuang, Xuanyu, et al.
Pubblicazione: (2024)
di: Zhuang, Xuanyu, et al.
Pubblicazione: (2024)
Investigating the Design Space of Diffusion Models for Speech Enhancement
di: Gonzalez, Philippe, et al.
Pubblicazione: (2023)
di: Gonzalez, Philippe, et al.
Pubblicazione: (2023)
Diffusion-Based Speech Enhancement in Matched and Mismatched Conditions Using a Heun-Based Sampler
di: Gonzalez, Philippe, et al.
Pubblicazione: (2023)
di: Gonzalez, Philippe, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Unsupervised Composable Representations for Audio
di: Bindi, Giovanni, et al.
Pubblicazione: (2024) -
Long-form music generation with latent diffusion
di: Evans, Zach, et al.
Pubblicazione: (2024) -
Sentiment analysis in non-fixed length audios using a Fully Convolutional Neural Network
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024) -
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
di: Messina, Francisco, et al.
Pubblicazione: (2025) -
Transformation of audio embeddings into interpretable, concept-based representations
di: Zhang, Alice, et al.
Pubblicazione: (2025)