Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Hsiang-Cheng, Li, You-Jin, Chao, Rong, Tsao, Yu, Su, Borching, Chien, Shao-Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026)
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
von: Hussain, Tassadaq, et al.
Veröffentlicht: (2024)
von: Hussain, Tassadaq, et al.
Veröffentlicht: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Visual-Informed Speech Enhancement Using Attention-Based Beamforming
von: Liu, Chihyun, et al.
Veröffentlicht: (2026)
von: Liu, Chihyun, et al.
Veröffentlicht: (2026)
Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database
von: Yaffe, Danielle, et al.
Veröffentlicht: (2025)
von: Yaffe, Danielle, et al.
Veröffentlicht: (2025)
Universal Speech Enhancement with Regression and Generative Mamba
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids
von: Saleem, Nasir, et al.
Veröffentlicht: (2025)
von: Saleem, Nasir, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Audio-Visual Pretrained Models to Improve Vocoded Speech Intelligibility in Cochlear Implant Simulation
von: Lai, Richard Lee, et al.
Veröffentlicht: (2023)
von: Lai, Richard Lee, et al.
Veröffentlicht: (2023)
EffortNet: A Deep Learning Framework for Objective Assessment of Speech Enhancement Technologies Using EEG-Based Alpha Oscillations
von: Sung, Ching-Chih, et al.
Veröffentlicht: (2025)
von: Sung, Ching-Chih, et al.
Veröffentlicht: (2025)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding
von: Shao, Mingchen, et al.
Veröffentlicht: (2026)
von: Shao, Mingchen, et al.
Veröffentlicht: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
von: Gelderblom, Femke B., et al.
Veröffentlicht: (2024)
von: Gelderblom, Femke B., et al.
Veröffentlicht: (2024)
From Evaluation to Optimization: Neural Speech Assessment for Downstream Applications
von: Tsao, Yu
Veröffentlicht: (2025)
von: Tsao, Yu
Veröffentlicht: (2025)
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2025)
Attention-Guided Adaptation for Code-Switching Speech Recognition
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
von: Aditya, Bobbi, et al.
Veröffentlicht: (2023)
French Listening Tests for the Assessment of Intelligibility, Quality, and Identity of Body-Conducted Speech Enhancement
von: Joubaud, Thomas, et al.
Veröffentlicht: (2025)
von: Joubaud, Thomas, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
An Investigation on Combining Geometry and Consistency Constraints into Phase Estimation for Speech Enhancement
von: Ho, Chun-Wei, et al.
Veröffentlicht: (2025)
von: Ho, Chun-Wei, et al.
Veröffentlicht: (2025)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2025)
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
Linguistic Knowledge Transfer Learning for Speech Enhancement
von: Hung, Kuo-Hsuan, et al.
Veröffentlicht: (2025)
von: Hung, Kuo-Hsuan, et al.
Veröffentlicht: (2025)
Dynamically Slimmable Speech Enhancement Network with Metric-Guided Training
von: Zhao, Haixin, et al.
Veröffentlicht: (2025)
von: Zhao, Haixin, et al.
Veröffentlicht: (2025)
AMDM-SE: Attention-based Multichannel Diffusion Model for Speech Enhancement
von: Opochinsky, Renana, et al.
Veröffentlicht: (2026)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2026)
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2024)
Improving the Robustness and Clinical Applicability of Automatic Respiratory Sound Classification Using Deep Learning-Based Audio Enhancement: Algorithm Development and Validation
von: Tzeng, Jing-Tong, et al.
Veröffentlicht: (2024)
von: Tzeng, Jing-Tong, et al.
Veröffentlicht: (2024)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
Learning Time-Graph Frequency Representation for Monaural Speech Enhancement
von: Wang, Tingting, et al.
Veröffentlicht: (2025)
von: Wang, Tingting, et al.
Veröffentlicht: (2025)
PLDNet: PLD-Guided Lightweight Deep Network Boosted by Efficient Attention for Handheld Dual-Microphone Speech Enhancement
von: Zhou, Nan, et al.
Veröffentlicht: (2024)
von: Zhou, Nan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026) -
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025) -
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
von: Ren, Wenze, et al.
Veröffentlicht: (2024) -
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025) -
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
von: Hussain, Tassadaq, et al.
Veröffentlicht: (2024)