RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Seongho, Choi, Yong-Hoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025)
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024)
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023)
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024)
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024)
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
von: Ko, Myeongjin, et al.
Veröffentlicht: (2023)
von: Ko, Myeongjin, et al.
Veröffentlicht: (2023)
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
von: Hono, Yukiya, et al.
Veröffentlicht: (2024)
von: Hono, Yukiya, et al.
Veröffentlicht: (2024)
Neural Vocoders as Speech Enhancers
von: Li, Andong, et al.
Veröffentlicht: (2025)
von: Li, Andong, et al.
Veröffentlicht: (2025)
BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Xiao-Hang, et al.
Veröffentlicht: (2024)
Vocoder-Projected Feature Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2025)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
Classification of Short Segment Pediatric Heart Sounds Based on a Transformer-Based Convolutional Neural Network
von: Hassanuzzaman, Md, et al.
Veröffentlicht: (2024)
von: Hassanuzzaman, Md, et al.
Veröffentlicht: (2024)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
Audio Classification of Low Feature Spectrograms Utilizing Convolutional Neural Networks
von: Elias, Noel
Veröffentlicht: (2024)
von: Elias, Noel
Veröffentlicht: (2024)
Automatic Equalization for Individual Instrument Tracks Using Convolutional Neural Networks
von: Mockenhaupt, Florian, et al.
Veröffentlicht: (2024)
von: Mockenhaupt, Florian, et al.
Veröffentlicht: (2024)
Barwise Section Boundary Detection in Symbolic Music Using Convolutional Neural Networks
von: Eldeeb, Omar, et al.
Veröffentlicht: (2025)
von: Eldeeb, Omar, et al.
Veröffentlicht: (2025)
Multi-stream Convolutional Neural Network with Frequency Selection for Robust Speaker Verification
von: Yao, Wei, et al.
Veröffentlicht: (2020)
von: Yao, Wei, et al.
Veröffentlicht: (2020)
Convolutional Neural Network Achieves Human-level Accuracy in Music Genre Classification
von: Dong, Mingwen
Veröffentlicht: (2018)
von: Dong, Mingwen
Veröffentlicht: (2018)
Sentiment analysis in non-fixed length audios using a Fully Convolutional Neural Network
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
InterGridNet: An Electric Network Frequency Approach for Audio Source Location Classification Using Convolutional Neural Networks
von: Korgialas, Christos, et al.
Veröffentlicht: (2025)
von: Korgialas, Christos, et al.
Veröffentlicht: (2025)
Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
MiSTR: Multi-Modal iEEG-to-Speech Synthesis with Transformer-Based Prosody Prediction and Neural Phase Reconstruction
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
von: Al-Radhi, Mohammed Salah, et al.
Veröffentlicht: (2025)
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
Temporal Convolution-based Hybrid Model Approach with Representation Learning for Real-Time Acoustic Anomaly Detection
von: Dissanayaka, Sahan, et al.
Veröffentlicht: (2024)
von: Dissanayaka, Sahan, et al.
Veröffentlicht: (2024)
Fine-grained Soundscape Control for Augmented Hearing
von: Oh, Seunghyun, et al.
Veröffentlicht: (2026)
von: Oh, Seunghyun, et al.
Veröffentlicht: (2026)
Adversarial Data Augmentation for Robust Speaker Verification
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenyu, et al.
Veröffentlicht: (2024)
Targeted Augmented Data for Audio Deepfake Detection
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Generating Music with Structure Using Self-Similarity as Attention
von: Hager, Sophia, et al.
Veröffentlicht: (2024)
von: Hager, Sophia, et al.
Veröffentlicht: (2024)
Conditional Generative Data Augmentation for Clinical Audio Datasets
von: Seibold, Matthias, et al.
Veröffentlicht: (2022)
von: Seibold, Matthias, et al.
Veröffentlicht: (2022)
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
Anticipatory Music Transformer
von: Thickstun, John, et al.
Veröffentlicht: (2023)
von: Thickstun, John, et al.
Veröffentlicht: (2023)
Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
von: Roman, Iran R., et al.
Veröffentlicht: (2024)
von: Roman, Iran R., et al.
Veröffentlicht: (2024)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
von: Fichtinger, Alexander, et al.
Veröffentlicht: (2025)
von: Fichtinger, Alexander, et al.
Veröffentlicht: (2025)
Combolutional Neural Networks
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
von: Ellinas, Nikolaos, et al.
Veröffentlicht: (2025) -
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis
von: Agrawal, Prabhav, et al.
Veröffentlicht: (2024) -
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
von: Shibuya, Takashi, et al.
Veröffentlicht: (2023) -
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems
von: Liu, Jeongmin, et al.
Veröffentlicht: (2024) -
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator
von: Kaneko, Takuhiro, et al.
Veröffentlicht: (2024)