Fast-ULCNet: A fast and ultra low complexity network for single-channel speech enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Larraza, Nicolás Arrieta, de Koeijer, Niels |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A two-step approach for speech enhancement in low-SNR scenarios using cyclostationary beamforming and DNNs
von: Bologni, Giovanni, et al.
Veröffentlicht: (2026)
von: Bologni, Giovanni, et al.
Veröffentlicht: (2026)
FINALLY: fast and universal speech enhancement with studio-like quality
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024)
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
von: Chi, Cheng, et al.
Veröffentlicht: (2024)
Automated evaluation of children's speech fluency for low-resource languages
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Online neural fusion of distortionless differential beamformers for robust speech enhancement
von: Qian, Yuanhang, et al.
Veröffentlicht: (2025)
von: Qian, Yuanhang, et al.
Veröffentlicht: (2025)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
von: Vecino, Biel Tura, et al.
Veröffentlicht: (2025)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
von: Gong, Rong, et al.
Veröffentlicht: (2024)
von: Gong, Rong, et al.
Veröffentlicht: (2024)
Deep low-latency joint speech transmission and enhancement over a gaussian channel
von: Bokaei, Mohammad, et al.
Veröffentlicht: (2024)
von: Bokaei, Mohammad, et al.
Veröffentlicht: (2024)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
Ensemble of classifiers for speech evaluation
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
von: Belokrylov, G., et al.
Veröffentlicht: (2024)
A correlation-permutation approach for speech-music encoders model merging
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Heterogeneous bimodal attention fusion for speech emotion recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
Inter-channel Conv-TasNet for multichannel speech enhancement
von: Lee, Dongheon, et al.
Veröffentlicht: (2021)
von: Lee, Dongheon, et al.
Veröffentlicht: (2021)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
Selective Classifier-free Guidance for Zero-shot Text-to-speech
von: Zheng, John, et al.
Veröffentlicht: (2025)
von: Zheng, John, et al.
Veröffentlicht: (2025)
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025)
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025)
SPMamba: State-space model is all you need in speech separation
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Enhancing dysarthria speech feature representation with empirical mode decomposition and Walsh-Hadamard transform
von: Zhu, Ting, et al.
Veröffentlicht: (2023)
von: Zhu, Ting, et al.
Veröffentlicht: (2023)
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
learning discriminative features from spectrograms using center loss for speech emotion recognition
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
von: Dai, Dongyang, et al.
Veröffentlicht: (2025)
Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks
von: Beguš, Gašper, et al.
Veröffentlicht: (2023)
von: Beguš, Gašper, et al.
Veröffentlicht: (2023)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
Tone recognition in low-resource languages of North-East India: peeling the layers of SSL-based speech models
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
von: Gogoi, Parismita, et al.
Veröffentlicht: (2025)
Perceptual implications of automatic anonymization in pathological speech
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2025)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2025)
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
von: Samui, Suman
Veröffentlicht: (2022)
von: Samui, Suman
Veröffentlicht: (2022)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2026)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2026)
Index-MSR: A high-efficiency multimodal fusion framework for speech recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
von: Chen, Jinming, et al.
Veröffentlicht: (2025)
Spectrogram features for audio and speech analysis
von: McLoughlin, Ian, et al.
Veröffentlicht: (2026)
von: McLoughlin, Ian, et al.
Veröffentlicht: (2026)
Enhancing nonnative speech perception and production through an AI-powered application
von: Georgiou, Georgios P.
Veröffentlicht: (2025)
von: Georgiou, Georgios P.
Veröffentlicht: (2025)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
Single-channel speech enhancement using learnable loss mixup
von: Chang, Oscar, et al.
Veröffentlicht: (2023)
von: Chang, Oscar, et al.
Veröffentlicht: (2023)
Incremental FastPitch: Chunk-based High Quality Text to Speech
von: Du, Muyang, et al.
Veröffentlicht: (2024)
von: Du, Muyang, et al.
Veröffentlicht: (2024)
A unified front-end framework for English text-to-speech synthesis
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
von: Ying, Zelin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A two-step approach for speech enhancement in low-SNR scenarios using cyclostationary beamforming and DNNs
von: Bologni, Giovanni, et al.
Veröffentlicht: (2026) -
FINALLY: fast and universal speech enhancement with studio-like quality
von: Babaev, Nicholas, et al.
Veröffentlicht: (2024) -
End-to-end multi-channel speaker extraction and binaural speech synthesis
von: Chi, Cheng, et al.
Veröffentlicht: (2024) -
Automated evaluation of children's speech fluency for low-resource languages
von: Zhang, Bowen, et al.
Veröffentlicht: (2025) -
Online neural fusion of distortionless differential beamformers for robust speech enhancement
von: Qian, Yuanhang, et al.
Veröffentlicht: (2025)