Audio-Visual Speech Enhancement: Architectural Design and Deployment Strategies
Fuente:
arXiv
Saved in:
| Main Authors: | Hamadouche, Anis, Luo, Haifeng, Sellathurai, Mathini, Hussain, Amir, Ratnarajah, Tharm |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Antenna Health-Aware Selective Beamforming for Hardware-Constrained DFRC Systems I
by: Hamadouche, Anis, et al.
Published: (2024)
by: Hamadouche, Anis, et al.
Published: (2024)
Antenna Health-Aware Selective Beamforming for Hardware-Constrained DFRC Systems II
by: Hamadouche, Anis, et al.
Published: (2024)
by: Hamadouche, Anis, et al.
Published: (2024)
Efficient Dual-Blind Deconvolution for Joint Radar-Communication Systems Using ADMM: Enhancing Channel Estimation and Signal Recovery in 5G mmWave Networks
by: Hamadouche, Anis, et al.
Published: (2024)
by: Hamadouche, Anis, et al.
Published: (2024)
Context-Enhanced CSI Tracking Using Koopman-Inspired Dual Autoencoders in Dynamic Wireless Environments
by: Hamadouche, Anis, et al.
Published: (2024)
by: Hamadouche, Anis, et al.
Published: (2024)
Trade-offs in Reliability and Performance Using Selective Beamforming for Ultra-Massive MIMO
by: Hamadouche, Anis, et al.
Published: (2024)
by: Hamadouche, Anis, et al.
Published: (2024)
Leveraging Kernel Symmetry for Joint Compression and Error Mitigation in Edge Model Transfer
by: Hamadouche, Anis, et al.
Published: (2026)
by: Hamadouche, Anis, et al.
Published: (2026)
Reconfigurable FPGA-Based Solvers For Sparse Satellite Control
by: Hamadouche, Anis, et al.
Published: (2024)
by: Hamadouche, Anis, et al.
Published: (2024)
A Study on Speech Assessment with Visual Cues
by: Ahmed, Shafique, et al.
Published: (2025)
by: Ahmed, Shafique, et al.
Published: (2025)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
by: Hussain, Tassadaq, et al.
Published: (2024)
by: Hussain, Tassadaq, et al.
Published: (2024)
Toward Universal Speech Enhancement for Diverse Input Conditions
by: Zhang, Wangyou, et al.
Published: (2023)
by: Zhang, Wangyou, et al.
Published: (2023)
Significance of Chirp MFCC as a Feature in Speech and Audio Applications
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
by: Wang, Kuan-Chen, et al.
Published: (2024)
by: Wang, Kuan-Chen, et al.
Published: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
by: Zhang, Wangyou, et al.
Published: (2025)
by: Zhang, Wangyou, et al.
Published: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
by: Sato, Hiroshi, et al.
Published: (2025)
by: Sato, Hiroshi, et al.
Published: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
by: Serre, Thomas, et al.
Published: (2026)
by: Serre, Thomas, et al.
Published: (2026)
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
by: Ratnarajah, Anton, et al.
Published: (2026)
by: Ratnarajah, Anton, et al.
Published: (2026)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
by: Masuyama, Yoshiki, et al.
Published: (2025)
by: Masuyama, Yoshiki, et al.
Published: (2025)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
by: Wang, Ziqian, et al.
Published: (2025)
by: Wang, Ziqian, et al.
Published: (2025)
Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
by: Koo, Inyong, et al.
Published: (2026)
by: Koo, Inyong, et al.
Published: (2026)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
by: Shetu, Shrishti Saha, et al.
Published: (2024)
by: Shetu, Shrishti Saha, et al.
Published: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
by: Yang, Yujie, et al.
Published: (2025)
by: Yang, Yujie, et al.
Published: (2025)
Speech Enhancement Based on Drifting Models
by: Xu, Liang, et al.
Published: (2026)
by: Xu, Liang, et al.
Published: (2026)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
by: Bae, Hanbin, et al.
Published: (2024)
by: Bae, Hanbin, et al.
Published: (2024)
QINCODEC: Neural Audio Compression with Implicit Neural Codebooks
by: Lahrichi, Zineb, et al.
Published: (2025)
by: Lahrichi, Zineb, et al.
Published: (2025)
Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
by: Carson, Alistair, et al.
Published: (2024)
by: Carson, Alistair, et al.
Published: (2024)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
by: Yan, Haoyin, et al.
Published: (2024)
by: Yan, Haoyin, et al.
Published: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
by: Qi, Tianhua, et al.
Published: (2026)
by: Qi, Tianhua, et al.
Published: (2026)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
by: Hao, Xiang, et al.
Published: (2020)
by: Hao, Xiang, et al.
Published: (2020)
BANC: Towards Efficient Binaural Audio Neural Codec for Overlapping Speech
by: Ratnarajah, Anton, et al.
Published: (2023)
by: Ratnarajah, Anton, et al.
Published: (2023)
Diffusion-based Unsupervised Audio-visual Speech Enhancement
by: Ayilo, Jean-Eudes, et al.
Published: (2024)
by: Ayilo, Jean-Eudes, et al.
Published: (2024)
A Data-Centric Approach to Generalizable Speech Deepfake Detection
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
by: Xu, Zhongweiyang, et al.
Published: (2024)
by: Xu, Zhongweiyang, et al.
Published: (2024)
Modulation Feature Enhancement with a Multi-Stage Attention Network for Underwater Acoustic Target Recognition
by: Yu, Jiaping, et al.
Published: (2026)
by: Yu, Jiaping, et al.
Published: (2026)
A Multi-decoder Neural Tracking Method for Accurately Predicting Speech Intelligibility
by: Sonck, Rien, et al.
Published: (2026)
by: Sonck, Rien, et al.
Published: (2026)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
by: Kim, Minje, et al.
Published: (2024)
by: Kim, Minje, et al.
Published: (2024)
DECAF: Dynamic Envelope Context-Aware Fusion for Speech-Envelope Reconstruction from EEG
by: Thakkar, Karan, et al.
Published: (2026)
by: Thakkar, Karan, et al.
Published: (2026)
ASVspoof 5: Evaluation of Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
by: Tuncay, Ludovic, et al.
Published: (2025)
by: Tuncay, Ludovic, et al.
Published: (2025)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
by: Ronchini, Francesca, et al.
Published: (2025)
by: Ronchini, Francesca, et al.
Published: (2025)
Similar Items
-
Antenna Health-Aware Selective Beamforming for Hardware-Constrained DFRC Systems I
by: Hamadouche, Anis, et al.
Published: (2024) -
Antenna Health-Aware Selective Beamforming for Hardware-Constrained DFRC Systems II
by: Hamadouche, Anis, et al.
Published: (2024) -
Efficient Dual-Blind Deconvolution for Joint Radar-Communication Systems Using ADMM: Enhancing Channel Estimation and Signal Recovery in 5G mmWave Networks
by: Hamadouche, Anis, et al.
Published: (2024) -
Context-Enhanced CSI Tracking Using Koopman-Inspired Dual Autoencoders in Dynamic Wireless Environments
by: Hamadouche, Anis, et al.
Published: (2024) -
Trade-offs in Reliability and Performance Using Selective Beamforming for Ultra-Massive MIMO
by: Hamadouche, Anis, et al.
Published: (2024)