WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baoueb, Teysir, Bie, Xiaoyu, Janati, Hicham, Richard, Gael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
von: Li, Haowen, et al.
Veröffentlicht: (2026)
von: Li, Haowen, et al.
Veröffentlicht: (2026)
Désentrelacement Fréquentiel Doux pour les Codecs Audio Neuronaux
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
Learning Source Disentanglement in Neural Audio Codec
von: Bie, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Bie, Xiaoyu, et al.
Veröffentlicht: (2024)
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
von: Mancusi, Michele, et al.
Veröffentlicht: (2024)
von: Mancusi, Michele, et al.
Veröffentlicht: (2024)
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
von: Raj, Desh
Veröffentlicht: (2024)
von: Raj, Desh
Veröffentlicht: (2024)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
Timbre Difference Capturing in Anomalous Sound Detection
von: Nishida, Tomoya, et al.
Veröffentlicht: (2024)
von: Nishida, Tomoya, et al.
Veröffentlicht: (2024)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
Timbre Perception, Representation, and its Neuroscientific Exploration: A Comprehensive Review
von: Zhang, Hong, et al.
Veröffentlicht: (2024)
von: Zhang, Hong, et al.
Veröffentlicht: (2024)
DIFFRENT: A Diffusion Model for Recording Environment Transfer of Speech
von: Im, Jaekwon, et al.
Veröffentlicht: (2024)
von: Im, Jaekwon, et al.
Veröffentlicht: (2024)
Music Style Transfer with Time-Varying Inversion of Diffusion Models
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
von: Wu, Zhiyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2025)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
Is Transfer Learning Necessary for Violin Transcription?
von: Peng, Yueh-Po, et al.
Veröffentlicht: (2025)
von: Peng, Yueh-Po, et al.
Veröffentlicht: (2025)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
Diff-MST: Differentiable Mixing Style Transfer
von: Vanka, Soumya Sai, et al.
Veröffentlicht: (2024)
von: Vanka, Soumya Sai, et al.
Veröffentlicht: (2024)
Transferable Adversarial Attacks on Audio Deepfake Detection
von: Farooq, Muhammad Umar, et al.
Veröffentlicht: (2025)
von: Farooq, Muhammad Umar, et al.
Veröffentlicht: (2025)
Diff-MSTC: A Mixing Style Transfer Prototype for Cubase
von: Vanka, Soumya Sai, et al.
Veröffentlicht: (2024)
von: Vanka, Soumya Sai, et al.
Veröffentlicht: (2024)
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
von: Li, Jingbei, et al.
Veröffentlicht: (2023)
von: Li, Jingbei, et al.
Veröffentlicht: (2023)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
Relative Transfer Matrix Estimator using Covariance Subtraction
von: Manamperi, Wageesha N., et al.
Veröffentlicht: (2025)
von: Manamperi, Wageesha N., et al.
Veröffentlicht: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2024)
End-to-End Multi-Task Learning for Adjustable Joint Noise Reduction and Hearing Loss Compensation
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2026)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2026)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
Music2Fail: Transfer Music to Failed Recorder Style
von: Leong, Chon In, et al.
Veröffentlicht: (2024)
von: Leong, Chon In, et al.
Veröffentlicht: (2024)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025) -
Diff-TONE: Timestep Optimization for iNstrument Editing in Text-to-Music Diffusion Models
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025) -
GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model
von: Liu, Haocheng, et al.
Veröffentlicht: (2024) -
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024) -
Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)