SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baoueb, Teysir, Liu, Haocheng, Fontaine, Mathieu, Roux, Jonathan Le, Richard, Gael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
Speech dereverberation constrained on room impulse response characteristics
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
von: Bahrman, Louis, et al.
Veröffentlicht: (2025)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
DiffAU: Diffusion-Based Ambisonics Upscaling
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
Physics-Informed Direction-Aware Neural Acoustic Fields
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2026)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2026)
SIRUP: A diffusion-based virtual upmixer of steering vectors for highly-directive spatialization with first-order ambisonics
von: Picard, Emilio, et al.
Veröffentlicht: (2026)
von: Picard, Emilio, et al.
Veröffentlicht: (2026)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
FasTUSS: Faster Task-Aware Unified Source Separation
von: Paissan, Francesco, et al.
Veröffentlicht: (2025)
von: Paissan, Francesco, et al.
Veröffentlicht: (2025)
Modèle physique variationnel pour l'estimation de réponses impulsionnelles de salles
von: Lalay, Louis, et al.
Veröffentlicht: (2025)
von: Lalay, Louis, et al.
Veröffentlicht: (2025)
Online speaker diarization of meetings guided by speech separation
von: Gruttadauria, Elio, et al.
Veröffentlicht: (2024)
von: Gruttadauria, Elio, et al.
Veröffentlicht: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Factorized RVQ-GAN For Disentangled Speech Tokenization
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
Unsupervised Harmonic Parameter Estimation Using Differentiable DSP and Spectral Optimal Transport
von: Torres, Bernardo, et al.
Veröffentlicht: (2023)
von: Torres, Bernardo, et al.
Veröffentlicht: (2023)
Musical Source Separation of Brazilian Percussion
von: Namballa, Richa, et al.
Veröffentlicht: (2025)
von: Namballa, Richa, et al.
Veröffentlicht: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
AutoMashup: Automatic Music Mashups Creation
von: Delabaere, Marine, et al.
Veröffentlicht: (2025)
von: Delabaere, Marine, et al.
Veröffentlicht: (2025)
Musical Score Following using Statistical Inference
von: Cowley, Josephine
Veröffentlicht: (2025)
von: Cowley, Josephine
Veröffentlicht: (2025)
How Does Instrumental Music Help SingFake Detection?
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
The Inverse Drum Machine: Source Separation Through Joint Transcription and Analysis-by-Synthesis
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
von: Torres, Bernardo, et al.
Veröffentlicht: (2025)
Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning
von: Lee, Sungho, et al.
Veröffentlicht: (2025)
von: Lee, Sungho, et al.
Veröffentlicht: (2025)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Note-Level Singing Melody Transcription for Time-Aligned Musical Score Generation
von: Kim, Leekyung, et al.
Veröffentlicht: (2025)
von: Kim, Leekyung, et al.
Veröffentlicht: (2025)
Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
von: Shore, Noah
Veröffentlicht: (2025)
von: Shore, Noah
Veröffentlicht: (2025)
Do Music Source Separation Models Preserve Spatial Information in Binaural Audio?
von: Namballa, Richa, et al.
Veröffentlicht: (2025)
von: Namballa, Richa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model
von: Liu, Haocheng, et al.
Veröffentlicht: (2024) -
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025) -
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025) -
Speech dereverberation constrained on room impulse response characteristics
von: Bahrman, Louis, et al.
Veröffentlicht: (2024) -
WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)