BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Hui-Peng, Lu, Ye-Xin, Ai, Yang, Ling, Zhen-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
von: Shi, Yu-Fei, et al.
Veröffentlicht: (2024)
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Vision-Integrated High-Quality Neural Speech Coding
von: Guo, Yao, et al.
Veröffentlicht: (2025)
von: Guo, Yao, et al.
Veröffentlicht: (2025)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2023)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
von: Yang, Runxuan, et al.
Veröffentlicht: (2025)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
von: Ai, Yang, et al.
Veröffentlicht: (2024)
von: Ai, Yang, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
von: Hong, Seongho, et al.
Veröffentlicht: (2025)
Wave-Trainer-Fit: Neural Vocoder with Trainable Prior and Fixed-Point Iteration towards High-Quality Speech Generation from SSL features
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2026)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
Stage-Wise and Prior-Aware Neural Speech Phase Prediction
von: Liu, Fei, et al.
Veröffentlicht: (2024)
von: Liu, Fei, et al.
Veröffentlicht: (2024)
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024)
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
von: Hono, Yukiya, et al.
Veröffentlicht: (2024)
von: Hono, Yukiya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024) -
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024) -
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025) -
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025) -
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)