Semantic Communications for Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Weng, Zhenzi, Qin, Zhijin, Li, Geoffrey Ye |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2021
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robust Semantic Communications for Speech Transmission
von: Weng, Zhenzi, et al.
Veröffentlicht: (2024)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2024)
Large Model Empowered Streaming Speech Semantic Communications
von: Weng, Zhenzi, et al.
Veröffentlicht: (2025)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2025)
Semantic MIMO Systems for Speech-to-Text Transmission
von: Weng, Zhenzi, et al.
Veröffentlicht: (2024)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
A Study on Speech Assessment with Visual Cues
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Speech dereverberation constrained on room impulse response characteristics
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
von: Rahimi, Akam, et al.
Veröffentlicht: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
von: Kwon, Younghoo, et al.
Veröffentlicht: (2024)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
von: Wills, Simone, et al.
Veröffentlicht: (2023)
von: Wills, Simone, et al.
Veröffentlicht: (2023)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
Speech-preserving active noise control: a deep learning approach in reverberant environments
von: Dai, Shuning
Veröffentlicht: (2026)
von: Dai, Shuning
Veröffentlicht: (2026)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
von: Phan, Dang Thoai, et al.
Veröffentlicht: (2025)
von: Phan, Dang Thoai, et al.
Veröffentlicht: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Detecting Post-Stroke Aphasia Via Brain Responses to Speech in a Deep Learning Framework
von: De Clercq, Pieter, et al.
Veröffentlicht: (2024)
von: De Clercq, Pieter, et al.
Veröffentlicht: (2024)
Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
von: Tang, Chenyu, et al.
Veröffentlicht: (2023)
von: Tang, Chenyu, et al.
Veröffentlicht: (2023)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
Breaking Speaker Recognition with PaddingBack
von: Ye, Zhe, et al.
Veröffentlicht: (2023)
von: Ye, Zhe, et al.
Veröffentlicht: (2023)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
von: Kheddar, Hamza, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Robust Semantic Communications for Speech Transmission
von: Weng, Zhenzi, et al.
Veröffentlicht: (2024) -
Large Model Empowered Streaming Speech Semantic Communications
von: Weng, Zhenzi, et al.
Veröffentlicht: (2025) -
Semantic MIMO Systems for Speech-to-Text Transmission
von: Weng, Zhenzi, et al.
Veröffentlicht: (2024) -
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024) -
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)