Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Yongjoon, Kim, Chanwoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
di: Zhou, Xuanru, et al.
Pubblicazione: (2024)
di: Zhou, Xuanru, et al.
Pubblicazione: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
di: Yun, Jun-Hak, et al.
Pubblicazione: (2025)
di: Yun, Jun-Hak, et al.
Pubblicazione: (2025)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
di: Ma, Zhengrui, et al.
Pubblicazione: (2024)
di: Ma, Zhengrui, et al.
Pubblicazione: (2024)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
di: Chen, Jinming, et al.
Pubblicazione: (2024)
di: Chen, Jinming, et al.
Pubblicazione: (2024)
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
From Modular to End-to-End Speaker Diarization
di: Landini, Federico
Pubblicazione: (2024)
di: Landini, Federico
Pubblicazione: (2024)
An End-to-End Approach for Chord-Conditioned Song Generation
di: Gao, Shuochen, et al.
Pubblicazione: (2024)
di: Gao, Shuochen, et al.
Pubblicazione: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
di: Liu, Xiaoqian, et al.
Pubblicazione: (2024)
di: Liu, Xiaoqian, et al.
Pubblicazione: (2024)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
di: Singh, Prachi, et al.
Pubblicazione: (2024)
di: Singh, Prachi, et al.
Pubblicazione: (2024)
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
di: Wang, Junyu, et al.
Pubblicazione: (2024)
di: Wang, Junyu, et al.
Pubblicazione: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
di: Yuan, Hui-Guan, et al.
Pubblicazione: (2025)
di: Yuan, Hui-Guan, et al.
Pubblicazione: (2025)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
di: Alvarez-Trejos, Juan Ignacio, et al.
Pubblicazione: (2024)
di: Alvarez-Trejos, Juan Ignacio, et al.
Pubblicazione: (2024)
End-to-End User-Defined Keyword Spotting using Shifted Delta Coefficients
di: V, Kesavaraj, et al.
Pubblicazione: (2024)
di: V, Kesavaraj, et al.
Pubblicazione: (2024)
A High-Fidelity Speech Super Resolution Network using a Complex Global Attention Module with Spectro-Temporal Loss
di: Tamiti, Tarikul Islam, et al.
Pubblicazione: (2025)
di: Tamiti, Tarikul Islam, et al.
Pubblicazione: (2025)
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
di: Ryu, Yerin, et al.
Pubblicazione: (2025)
di: Ryu, Yerin, et al.
Pubblicazione: (2025)
End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding
di: Zeng, Wei, et al.
Pubblicazione: (2024)
di: Zeng, Wei, et al.
Pubblicazione: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
di: Polák, Peter, et al.
Pubblicazione: (2023)
di: Polák, Peter, et al.
Pubblicazione: (2023)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
di: He, Xinlu, et al.
Pubblicazione: (2025)
di: He, Xinlu, et al.
Pubblicazione: (2025)
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
di: Jinghua, Liang, et al.
Pubblicazione: (2026)
di: Jinghua, Liang, et al.
Pubblicazione: (2026)
LearnAFE: Circuit-Algorithm Co-design Framework for Learnable Audio Analog Front-End
di: Hu, Jinhai, et al.
Pubblicazione: (2025)
di: Hu, Jinhai, et al.
Pubblicazione: (2025)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
di: Ok, Hyunjong, et al.
Pubblicazione: (2025)
di: Ok, Hyunjong, et al.
Pubblicazione: (2025)
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
di: Min, Anna, et al.
Pubblicazione: (2025)
di: Min, Anna, et al.
Pubblicazione: (2025)
TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms
di: Sui, Yueyuan, et al.
Pubblicazione: (2024)
di: Sui, Yueyuan, et al.
Pubblicazione: (2024)
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
di: Tan, Wei, et al.
Pubblicazione: (2025)
di: Tan, Wei, et al.
Pubblicazione: (2025)
Incremental FastPitch: Chunk-based High Quality Text to Speech
di: Du, Muyang, et al.
Pubblicazione: (2024)
di: Du, Muyang, et al.
Pubblicazione: (2024)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
di: Zhou, Xuanru, et al.
Pubblicazione: (2024)
di: Zhou, Xuanru, et al.
Pubblicazione: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)
di: Chi, Cheng, et al.
Pubblicazione: (2024)
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
di: Kühne, Nikolai Lund, et al.
Pubblicazione: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
di: Guimarães, Heitor R., et al.
Pubblicazione: (2024)
di: Guimarães, Heitor R., et al.
Pubblicazione: (2024)
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
di: Lee, Junseok, et al.
Pubblicazione: (2026)
di: Lee, Junseok, et al.
Pubblicazione: (2026)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
di: Vecino, Biel Tura, et al.
Pubblicazione: (2025)
di: Vecino, Biel Tura, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
di: Zhou, Xuanru, et al.
Pubblicazione: (2024) -
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024) -
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
di: Yun, Jun-Hak, et al.
Pubblicazione: (2025) -
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
di: Ma, Zhengrui, et al.
Pubblicazione: (2024) -
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
di: Chen, Jinming, et al.
Pubblicazione: (2024)