Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiaoyu, Li, Xu, Serrà, Joan, Pascual, Santiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
Semantic Communications for Speech Recognition
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
von: Weng, Zhenzi, et al.
Veröffentlicht: (2021)
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
von: Baoueb, Teysir, et al.
Veröffentlicht: (2025)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
Differentiable Acoustic Radiance Transfer
von: Lee, Sungho, et al.
Veröffentlicht: (2025)
von: Lee, Sungho, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
von: Yao, Shengshi, et al.
Veröffentlicht: (2025)
von: Yao, Shengshi, et al.
Veröffentlicht: (2025)
Decomposing the Influence of Physical Acoustic Modeling on Neural Personal Sound Zone Rendering: An Ablation Study
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
von: Jiang, Hao, et al.
Veröffentlicht: (2026)
Physics-Informed Direction-Aware Neural Acoustic Fields
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2026)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2026)
Acoustical Features as Knee Health Biomarkers: A Critical Analysis
von: Kechris, Christodoulos, et al.
Veröffentlicht: (2024)
von: Kechris, Christodoulos, et al.
Veröffentlicht: (2024)
Machine Learning in Acoustics: A Review and Open-Source Repository
von: McCarthy, Ryan A., et al.
Veröffentlicht: (2025)
von: McCarthy, Ryan A., et al.
Veröffentlicht: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
von: Yeon, Inmo, et al.
Veröffentlicht: (2023)
von: Yeon, Inmo, et al.
Veröffentlicht: (2023)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
von: Fan, Junyi, et al.
Veröffentlicht: (2025)
Comparison of Classification Algorithms for COVID19 Detection using Cough Acoustic Signals
von: Erdoğan, Yunus Emre, et al.
Veröffentlicht: (2022)
von: Erdoğan, Yunus Emre, et al.
Veröffentlicht: (2022)
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
von: Phan, Dang Thoai, et al.
Veröffentlicht: (2025)
von: Phan, Dang Thoai, et al.
Veröffentlicht: (2025)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
von: Chen, Xiaodan, et al.
Veröffentlicht: (2025)
RIR-Mega-Speech: A Reverberant Speech Corpus with Comprehensive Acoustic Metadata and Reproducible Evaluation
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
Acoustic Simulation Framework for Multi-channel Replay Speech Detection
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization
von: Goswami, Mandip
Veröffentlicht: (2026)
von: Goswami, Mandip
Veröffentlicht: (2026)
Dynamic Prediction of Full-Ocean Depth SSP by Hierarchical LSTM: An Experimental Result
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
Joint Source-Environment Adaptation of Data-Driven Underwater Acoustic Source Ranging Based on Model Uncertainty
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
von: Kari, Dariush, et al.
Veröffentlicht: (2025)
A Study on Speech Assessment with Visual Cues
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
von: Ahmed, Shafique, et al.
Veröffentlicht: (2025)
Lightweight and Generalizable Acoustic Scene Representations via Contrastive Fine-Tuning and Distillation
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
von: Yuan, Kuang, et al.
Veröffentlicht: (2025)
HiRIS: an Airborne Sonar Sensor with a 1024 Channel Microphone Array for In-Air Acoustic Imaging
von: Laurijssen, Dennis, et al.
Veröffentlicht: (2024)
von: Laurijssen, Dennis, et al.
Veröffentlicht: (2024)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Speech dereverberation constrained on room impulse response characteristics
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
von: Bahrman, Louis, et al.
Veröffentlicht: (2024)
Relating the Neural Representations of Vocalized, Mimed, and Imagined Speech
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
von: Maghsoudi, Maryam, et al.
Veröffentlicht: (2026)
Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
von: Tang, Chenyu, et al.
Veröffentlicht: (2023)
von: Tang, Chenyu, et al.
Veröffentlicht: (2023)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024) -
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
von: Gállego, Gerard I., et al.
Veröffentlicht: (2024) -
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026) -
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024) -
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
von: Hao, Xiang, et al.
Veröffentlicht: (2020)