Enregistré dans:
| Auteurs principaux: | Shen, Rubing, Ren, Yanzhen, Sun, Zongkun |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2407.04575 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
par: Du, Hui-Peng, et autres
Publié: (2025)
par: Du, Hui-Peng, et autres
Publié: (2025)
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
par: Xu, Nan, et autres
Publié: (2025)
par: Xu, Nan, et autres
Publié: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
par: Chen, Shaowen, et autres
Publié: (2025)
par: Chen, Shaowen, et autres
Publié: (2025)
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
par: Cao, Yubing, et autres
Publié: (2024)
par: Cao, Yubing, et autres
Publié: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
par: Shibuya, Takashi, et autres
Publié: (2023)
par: Shibuya, Takashi, et autres
Publié: (2023)
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
par: Zhou, Wangzixi, et autres
Publié: (2026)
par: Zhou, Wangzixi, et autres
Publié: (2026)
Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
par: Yao, Zengwei, et autres
Publié: (2025)
par: Yao, Zengwei, et autres
Publié: (2025)
Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN
par: Zhang, Shiqi, et autres
Publié: (2024)
par: Zhang, Shiqi, et autres
Publié: (2024)
Neural Vocoders as Speech Enhancers
par: Li, Andong, et autres
Publié: (2025)
par: Li, Andong, et autres
Publié: (2025)
A Distilled Low-Latency Neural Vocoder with Explicit Amplitude and Phase Prediction
par: Du, Hui-Peng, et autres
Publié: (2025)
par: Du, Hui-Peng, et autres
Publié: (2025)
GAN-Based Multi-Microphone Spatial Target Speaker Extraction
par: Shetu, Shrishti Saha, et autres
Publié: (2025)
par: Shetu, Shrishti Saha, et autres
Publié: (2025)
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
par: Serbest, Sanberk, et autres
Publié: (2025)
par: Serbest, Sanberk, et autres
Publié: (2025)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
par: Cho, Hyunjae, et autres
Publié: (2024)
par: Cho, Hyunjae, et autres
Publié: (2024)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
par: Nguyen, Tan Dat, et autres
Publié: (2024)
par: Nguyen, Tan Dat, et autres
Publié: (2024)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
par: Shetu, Shrishti Saha, et autres
Publié: (2025)
par: Shetu, Shrishti Saha, et autres
Publié: (2025)
MaskCycleGAN-based Whisper to Normal Speech Conversion
par: Gupta, K. Rohith, et autres
Publié: (2024)
par: Gupta, K. Rohith, et autres
Publié: (2024)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
par: Liu, Wuyang, et autres
Publié: (2023)
par: Liu, Wuyang, et autres
Publié: (2023)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
par: Du, Hui-Peng, et autres
Publié: (2024)
par: Du, Hui-Peng, et autres
Publié: (2024)
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
par: Baoueb, Teysir, et autres
Publié: (2024)
par: Baoueb, Teysir, et autres
Publié: (2024)
TLDiffGAN: A Latent Diffusion-GAN Framework with Temporal Information Fusion for Anomalous Sound Detection
par: Ma, Chengyuan, et autres
Publié: (2026)
par: Ma, Chengyuan, et autres
Publié: (2026)
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions
par: Du, Hui-Peng, et autres
Publié: (2024)
par: Du, Hui-Peng, et autres
Publié: (2024)
DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration
par: Tan, Shihong, et autres
Publié: (2026)
par: Tan, Shihong, et autres
Publié: (2026)
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
par: Gu, Yicheng, et autres
Publié: (2025)
par: Gu, Yicheng, et autres
Publié: (2025)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
par: Chung, HaeChun
Publié: (2025)
par: Chung, HaeChun
Publié: (2025)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
par: Gu, Yicheng, et autres
Publié: (2024)
par: Gu, Yicheng, et autres
Publié: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
par: Zhu, Ge, et autres
Publié: (2024)
par: Zhu, Ge, et autres
Publié: (2024)
Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
par: Valin, Jean-Marc, et autres
Publié: (2024)
par: Valin, Jean-Marc, et autres
Publié: (2024)
Factorized RVQ-GAN For Disentangled Speech Tokenization
par: Khurana, Sameer, et autres
Publié: (2025)
par: Khurana, Sameer, et autres
Publié: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
par: Baas, Matthew, et autres
Publié: (2023)
par: Baas, Matthew, et autres
Publié: (2023)
Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
par: Du, Hui-Peng, et autres
Publié: (2026)
par: Du, Hui-Peng, et autres
Publié: (2026)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
par: Shetu, Shrishti Saha, et autres
Publié: (2024)
par: Shetu, Shrishti Saha, et autres
Publié: (2024)
Towards Out-of-Distribution Detection in Vocoder Recognition via Latent Feature Reconstruction
par: Du, Renmingyue, et autres
Publié: (2024)
par: Du, Renmingyue, et autres
Publié: (2024)
Non-Causal to Causal SSL-Supported Transfer Learning: Towards a High-Performance Low-Latency Speech Vocoder
par: Shi, Renzheng, et autres
Publié: (2024)
par: Shi, Renzheng, et autres
Publié: (2024)
CMGAN: Conformer-based Metric GAN for Speech Enhancement
par: Cao, Ruizhe, et autres
Publié: (2022)
par: Cao, Ruizhe, et autres
Publié: (2022)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
par: Jiang, Xiao-Hang, et autres
Publié: (2024)
par: Jiang, Xiao-Hang, et autres
Publié: (2024)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
par: Yoneyama, Reo, et autres
Publié: (2025)
par: Yoneyama, Reo, et autres
Publié: (2025)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
par: Lv, Yuanjun, et autres
Publié: (2024)
par: Lv, Yuanjun, et autres
Publié: (2024)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
par: Lai, Richard Lee, et autres
Publié: (2023)
par: Lai, Richard Lee, et autres
Publié: (2023)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
par: Zhang, Tong, et autres
Publié: (2025)
par: Zhang, Tong, et autres
Publié: (2025)
Documents similaires
-
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
par: Du, Hui-Peng, et autres
Publié: (2025) -
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
par: Xu, Nan, et autres
Publié: (2025) -
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
par: Chen, Shaowen, et autres
Publié: (2025) -
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
par: Cao, Yubing, et autres
Publié: (2024) -
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
par: Shibuya, Takashi, et autres
Publié: (2023)