Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Valin, Jean-Marc, Mustafa, Ahmed, Büthe, Jan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DRED: Deep REDundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder
von: Valin, Jean-Marc, et al.
Veröffentlicht: (2022)
von: Valin, Jean-Marc, et al.
Veröffentlicht: (2022)
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
von: Büthe, Jan, et al.
Veröffentlicht: (2023)
von: Büthe, Jan, et al.
Veröffentlicht: (2023)
A lightweight and robust method for blind wideband-to-fullband extension of speech
von: Büthe, Jan, et al.
Veröffentlicht: (2024)
von: Büthe, Jan, et al.
Veröffentlicht: (2024)
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
von: Subramani, Krishna, et al.
Veröffentlicht: (2023)
von: Subramani, Krishna, et al.
Veröffentlicht: (2023)
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
von: Rowe, David, et al.
Veröffentlicht: (2025)
von: Rowe, David, et al.
Veröffentlicht: (2025)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2026)
von: Rautenberg, Frederik, et al.
Veröffentlicht: (2026)
Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
von: Togami, Masahito, et al.
Veröffentlicht: (2024)
von: Togami, Masahito, et al.
Veröffentlicht: (2024)
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
von: Wu, Chun Yat, et al.
Veröffentlicht: (2025)
von: Wu, Chun Yat, et al.
Veröffentlicht: (2025)
SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
von: Terashima, Ryo, et al.
Veröffentlicht: (2025)
Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
von: Gu, Yicheng, et al.
Veröffentlicht: (2025)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis
von: Wang, Huimeng, et al.
Veröffentlicht: (2026)
von: Wang, Huimeng, et al.
Veröffentlicht: (2026)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
A Low-Complexity Speech Codec Using Parametric Dithering for ASR
von: Murray, Ellison, et al.
Veröffentlicht: (2025)
von: Murray, Ellison, et al.
Veröffentlicht: (2025)
Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
von: Langman, Ryan, et al.
Veröffentlicht: (2024)
von: Langman, Ryan, et al.
Veröffentlicht: (2024)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
von: Lakshminarayana, Kishor Kayyar, et al.
Veröffentlicht: (2025)
von: Lakshminarayana, Kishor Kayyar, et al.
Veröffentlicht: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
von: Cho, Hyunjae, et al.
Veröffentlicht: (2024)
Leveraging Discriminative Latent Representations for Conditioning GAN-Based Speech Enhancement
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2025)
Direct Preference Optimization for Speech Autoregressive Diffusion Models
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
von: Liu, Zhijun, et al.
Veröffentlicht: (2025)
Periodicity Pitch Detection in Complex Harmonies on EEG Timeline Data
von: Heinze, Maria, et al.
Veröffentlicht: (2020)
von: Heinze, Maria, et al.
Veröffentlicht: (2020)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
von: Li, Bohan, et al.
Veröffentlicht: (2025)
von: Li, Bohan, et al.
Veröffentlicht: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
Hybrid Real- And Complex-Valued Neural Network Concept For Low-Complexity Phase-Aware Speech Enhancement
von: Fiorio, Luan Vinícius, et al.
Veröffentlicht: (2025)
von: Fiorio, Luan Vinícius, et al.
Veröffentlicht: (2025)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Decoding Order Matters in Autoregressive Speech Synthesis
von: Zhao, Minghui, et al.
Veröffentlicht: (2026)
von: Zhao, Minghui, et al.
Veröffentlicht: (2026)
AS-Speech: Adaptive Style For Speech Synthesis
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
von: Tokala, Vikas, et al.
Veröffentlicht: (2024)
von: Tokala, Vikas, et al.
Veröffentlicht: (2024)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
von: Zhang, En-Wei, et al.
Veröffentlicht: (2025)
von: Zhang, En-Wei, et al.
Veröffentlicht: (2025)
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2024)
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
von: Cai, Runyuan, et al.
Veröffentlicht: (2026)
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2025)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2025)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DRED: Deep REDundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder
von: Valin, Jean-Marc, et al.
Veröffentlicht: (2022) -
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
von: Büthe, Jan, et al.
Veröffentlicht: (2023) -
A lightweight and robust method for blind wideband-to-fullband extension of speech
von: Büthe, Jan, et al.
Veröffentlicht: (2024) -
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
von: Subramani, Krishna, et al.
Veröffentlicht: (2023) -
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
von: Rowe, David, et al.
Veröffentlicht: (2025)