WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Wangzixi, Okamoto, Takuma, Ohtani, Yamato, Sakti, Sakriani, Kawai, Hisashi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control
by: Zhou, Wangzixi, et al.
Published: (2026)
by: Zhou, Wangzixi, et al.
Published: (2026)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
by: Cooper, Erica, et al.
Published: (2025)
by: Cooper, Erica, et al.
Published: (2025)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
by: Heo, Hyun-Jun, et al.
Published: (2023)
by: Heo, Hyun-Jun, et al.
Published: (2023)
InceptionNeXt: When Inception Meets ConvNeXt
by: Yu, Weihao, et al.
Published: (2023)
by: Yu, Weihao, et al.
Published: (2023)
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
by: Lau, Kin Wai, et al.
Published: (2024)
by: Lau, Kin Wai, et al.
Published: (2024)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
by: Hirano, Yuta, et al.
Published: (2025)
by: Hirano, Yuta, et al.
Published: (2025)
EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition
by: Boudouri, Yassine El, et al.
Published: (2025)
by: Boudouri, Yassine El, et al.
Published: (2025)
Reviving ConvNeXt for Efficient Convolutional Diffusion Models
by: Kwon, Taesung, et al.
Published: (2026)
by: Kwon, Taesung, et al.
Published: (2026)
E-ConvNeXt: A Lightweight and Efficient ConvNeXt Variant with Cross-Stage Partial Connections
by: Wang, Fang, et al.
Published: (2025)
by: Wang, Fang, et al.
Published: (2025)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
by: Rossenbach, Nick, et al.
Published: (2024)
by: Rossenbach, Nick, et al.
Published: (2024)
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
by: Du, Hui-Peng, et al.
Published: (2025)
by: Du, Hui-Peng, et al.
Published: (2025)
Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification
by: Wu, Bin, et al.
Published: (2024)
by: Wu, Bin, et al.
Published: (2024)
FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder
by: Shen, Rubing, et al.
Published: (2024)
by: Shen, Rubing, et al.
Published: (2024)
Ensemble of radiomics and ConvNeXt for breast cancer diagnosis
by: Garza-Abdala, Jorge Alberto, et al.
Published: (2026)
by: Garza-Abdala, Jorge Alberto, et al.
Published: (2026)
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
by: Xu, Nan, et al.
Published: (2025)
by: Xu, Nan, et al.
Published: (2025)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
by: Adila, Aulia, et al.
Published: (2024)
by: Adila, Aulia, et al.
Published: (2024)
ConvNeXt with Histopathology-Specific Augmentations for Mitotic Figure Classification
by: Feki, Hana, et al.
Published: (2025)
by: Feki, Hana, et al.
Published: (2025)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
by: Chen, Shaowen, et al.
Published: (2025)
by: Chen, Shaowen, et al.
Published: (2025)
ConvXformer: Differentially Private Hybrid ConvNeXt-Transformer for Inertial Navigation
by: Tariq, Omer, et al.
Published: (2025)
by: Tariq, Omer, et al.
Published: (2025)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
by: Handoyo, Ahmad Alfani, et al.
Published: (2024)
by: Handoyo, Ahmad Alfani, et al.
Published: (2024)
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
by: Cao, Yubing, et al.
Published: (2024)
by: Cao, Yubing, et al.
Published: (2024)
Integrating ConvNeXt and Vision Transformers for Enhancing Facial Age Estimation
by: Maroun, Gaby, et al.
Published: (2025)
by: Maroun, Gaby, et al.
Published: (2025)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
by: Nguyen, Tan Dat, et al.
Published: (2024)
by: Nguyen, Tan Dat, et al.
Published: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
by: Shibuya, Takashi, et al.
Published: (2023)
by: Shibuya, Takashi, et al.
Published: (2023)
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
by: Tyndall, Geoffrey, et al.
Published: (2024)
by: Tyndall, Geoffrey, et al.
Published: (2024)
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
by: Meise, Adrian, et al.
Published: (2025)
by: Meise, Adrian, et al.
Published: (2025)
Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography
by: Molina-Román, Yusdivia, et al.
Published: (2025)
by: Molina-Román, Yusdivia, et al.
Published: (2025)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
by: Du, Hui-Peng, et al.
Published: (2024)
by: Du, Hui-Peng, et al.
Published: (2024)
GANeXt: A Fully ConvNeXt-Enhanced Generative Adversarial Network for MRI- and CBCT-to-CT Synthesis
by: Mei, Siyuan, et al.
Published: (2025)
by: Mei, Siyuan, et al.
Published: (2025)
Neural Vocoders as Speech Enhancers
by: Li, Andong, et al.
Published: (2025)
by: Li, Andong, et al.
Published: (2025)
EVCC: Enhanced Vision Transformer-ConvNeXt-CoAtNet Fusion for Classification
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
by: Hasan, Kazi Reyazul, et al.
Published: (2025)
Nondestructive testing of runny salted egg yolk based on improved ConvNeXt‐T
by: Haoran Chen, et al.
Published: (2024)
by: Haoran Chen, et al.
Published: (2024)
Slimmable ConvNeXt: Width-Adaptive Inference for Efficient Multi-Device Deployment
by: Haberer, Janek, et al.
Published: (2026)
by: Haberer, Janek, et al.
Published: (2026)
Mushroom image classification and recognition based on improved ConvNeXt V2
by: Shulong Zhang, et al.
Published: (2025)
by: Shulong Zhang, et al.
Published: (2025)
CNFA: ConvNeXt Fusion Attention Module for Age Recognition of the Tangerine Peel
by: Fuqin Deng, et al.
Published: (2024)
by: Fuqin Deng, et al.
Published: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
by: Zhu, Ge, et al.
Published: (2024)
by: Zhu, Ge, et al.
Published: (2024)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
by: Lai, Richard Lee, et al.
Published: (2023)
by: Lai, Richard Lee, et al.
Published: (2023)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
by: Yoneyama, Reo, et al.
Published: (2025)
by: Yoneyama, Reo, et al.
Published: (2025)
Towards HRTF Personalization using Denoising Diffusion Models
by: Sánchez, Juan Camilo Albarracín, et al.
Published: (2025)
by: Sánchez, Juan Camilo Albarracín, et al.
Published: (2025)
Similar Items
-
Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control
by: Zhou, Wangzixi, et al.
Published: (2026) -
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
by: Cooper, Erica, et al.
Published: (2025) -
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
by: Heo, Hyun-Jun, et al.
Published: (2023) -
InceptionNeXt: When Inception Meets ConvNeXt
by: Yu, Weihao, et al.
Published: (2023) -
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
by: Lau, Kin Wai, et al.
Published: (2024)