PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hono, Yukiya, Hashimoto, Kei, Nankaku, Yoshihiko, Tokuda, Keiichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
A Robust Method for Pitch Tracking in the Frequency Following Response using Harmonic Amplitude Summation Filterbank
von: Sadeghkhani, Sajad, et al.
Veröffentlicht: (2025)
von: Sadeghkhani, Sajad, et al.
Veröffentlicht: (2025)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Directional Selective Fixed-Filter Active Noise Control Based on a Convolutional Neural Network in Reverberant Environments
von: Wang, Boxiang, et al.
Veröffentlicht: (2026)
von: Wang, Boxiang, et al.
Veröffentlicht: (2026)
Cross-domain Neural Pitch and Periodicity Estimation
von: Morrison, Max, et al.
Veröffentlicht: (2023)
von: Morrison, Max, et al.
Veröffentlicht: (2023)
DiffAU: Diffusion-Based Ambisonics Upscaling
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
von: Milstein, Amit, et al.
Veröffentlicht: (2025)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Aliasing Reduction in Neural Amp Modeling by Smoothing Activations
von: Sato, Ryota, et al.
Veröffentlicht: (2025)
von: Sato, Ryota, et al.
Veröffentlicht: (2025)
Audio Compression using Periodic Gabor with Biorthogonal Exchange: Implementation Using the Zak Transform
von: Alimi, Roger, et al.
Veröffentlicht: (2025)
von: Alimi, Roger, et al.
Veröffentlicht: (2025)
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2026)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2026)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Using Neurogram Similarity Index Measure (NSIM) to Model Hearing Loss and Cochlear Neural Degeneration
von: Cheema, Ahsan J., et al.
Veröffentlicht: (2025)
von: Cheema, Ahsan J., et al.
Veröffentlicht: (2025)
Decomposing the Influence of Physical Acoustic Modeling on Neural Personal Sound Zone Rendering: An Ablation Study
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model
von: Delgado, Pablo M., et al.
Veröffentlicht: (2024)
von: Delgado, Pablo M., et al.
Veröffentlicht: (2024)
Aliasing-Free Neural Audio Synthesis
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
Physics-Informed Direction-Aware Neural Acoustic Fields
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Low-Complexity Neural Wind Noise Reduction for Audio Recordings
von: Eftekhari, Hesam, et al.
Veröffentlicht: (2025)
von: Eftekhari, Hesam, et al.
Veröffentlicht: (2025)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
Sample Rate Independent Recurrent Neural Networks for Audio Effects Processing
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
von: Carson, Alistair, et al.
Veröffentlicht: (2024)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
von: Qi, Tianhua, et al.
Veröffentlicht: (2025)
von: Qi, Tianhua, et al.
Veröffentlicht: (2025)
Real time fault detection in 3D printers using Convolutional Neural Networks and acoustic signals
von: Waheed, Muhammad Fasih, et al.
Veröffentlicht: (2026)
von: Waheed, Muhammad Fasih, et al.
Veröffentlicht: (2026)
Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
von: Wilroth, Johanna, et al.
Veröffentlicht: (2026)
von: Wilroth, Johanna, et al.
Veröffentlicht: (2026)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Future Full-Ocean Deep SSPs Prediction based on Hierarchical Long Short-Term Memory Neural Networks
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025) -
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024) -
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024) -
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025) -
GLA-Grad: A Griffin-Lim Extended Waveform Generation Diffusion Model
von: Liu, Haocheng, et al.
Veröffentlicht: (2024)