FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Rubing, Ren, Yanzhen, Sun, Zongkun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
von: Xu, Nan, et al.
Veröffentlicht: (2025)
von: Xu, Nan, et al.
Veröffentlicht: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
von: Chen, Shaowen, et al.
Veröffentlicht: (2025)
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
von: Zhou, Wangzixi, et al.
Veröffentlicht: (2026)
von: Zhou, Wangzixi, et al.
Veröffentlicht: (2026)
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
von: Yao, Zengwei, et al.
Veröffentlicht: (2025)
von: Yao, Zengwei, et al.
Veröffentlicht: (2025)
Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN
von: Zhang, Shiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Shiqi, et al.
Veröffentlicht: (2024)
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025)
von: Serbest, Sanberk, et al.
Veröffentlicht: (2025)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
MaskCycleGAN-based Whisper to Normal Speech Conversion
von: Gupta, K. Rohith, et al.
Veröffentlicht: (2024)
von: Gupta, K. Rohith, et al.
Veröffentlicht: (2024)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration
von: Tan, Shihong, et al.
Veröffentlicht: (2026)
von: Tan, Shihong, et al.
Veröffentlicht: (2026)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
von: Valin, Jean-Marc, et al.
Veröffentlicht: (2024)
von: Valin, Jean-Marc, et al.
Veröffentlicht: (2024)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
von: Chung, HaeChun
Veröffentlicht: (2025)
von: Chung, HaeChun
Veröffentlicht: (2025)
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
TLDiffGAN: A Latent Diffusion-GAN Framework with Temporal Information Fusion for Anomalous Sound Detection
von: Ma, Chengyuan, et al.
Veröffentlicht: (2026)
von: Ma, Chengyuan, et al.
Veröffentlicht: (2026)
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2026)
Non-Causal to Causal SSL-Supported Transfer Learning: Towards a High-Performance Low-Latency Speech Vocoder
von: Shi, Renzheng, et al.
Veröffentlicht: (2024)
von: Shi, Renzheng, et al.
Veröffentlicht: (2024)
Towards Out-of-Distribution Detection in Vocoder Recognition via Latent Feature Reconstruction
von: Du, Renmingyue, et al.
Veröffentlicht: (2024)
von: Du, Renmingyue, et al.
Veröffentlicht: (2024)
Factorized RVQ-GAN For Disentangled Speech Tokenization
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
von: Du, Chenpeng, et al.
Veröffentlicht: (2023)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
von: Lai, Richard Lee, et al.
Veröffentlicht: (2023)
von: Lai, Richard Lee, et al.
Veröffentlicht: (2023)
CMGAN: Conformer-based Metric GAN for Speech Enhancement
von: Cao, Ruizhe, et al.
Veröffentlicht: (2022)
von: Cao, Ruizhe, et al.
Veröffentlicht: (2022)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
von: Du, Hui-Peng, et al.
Veröffentlicht: (2025) -
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
von: Xu, Nan, et al.
Veröffentlicht: (2025) -
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
von: Chen, Shaowen, et al.
Veröffentlicht: (2025) -
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
von: Cao, Yubing, et al.
Veröffentlicht: (2024) -
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)