Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Nakazawa, Kazushi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
von: Zezario, Ryandhimas E.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E.
Veröffentlicht: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
von: Zhang, Qishan, et al.
Veröffentlicht: (2024)
von: Zhang, Qishan, et al.
Veröffentlicht: (2024)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
von: Cao, Boxuan, et al.
Veröffentlicht: (2025)
von: Cao, Boxuan, et al.
Veröffentlicht: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
Audio is all in one: speech-driven gesture synthetics using WavLM pre-trained model
von: Zhang, Fan, et al.
Veröffentlicht: (2023)
von: Zhang, Fan, et al.
Veröffentlicht: (2023)
WavLM model ensemble for audio deepfake detection
von: Combei, David, et al.
Veröffentlicht: (2024)
von: Combei, David, et al.
Veröffentlicht: (2024)
LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
von: Shekar, Ram C. M. C., et al.
Veröffentlicht: (2025)
von: Shekar, Ram C. M. C., et al.
Veröffentlicht: (2025)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
Audio Deepfake Detection with Self-Supervised WavLM and Multi-Fusion Attentive Classifier
von: Guo, Yinlin, et al.
Veröffentlicht: (2023)
von: Guo, Yinlin, et al.
Veröffentlicht: (2023)
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
von: Lin, Jingru, et al.
Veröffentlicht: (2024)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
Multi-Speaker DOA Estimation in Binaural Hearing Aids using Deep Learning and Speaker Count Fusion
von: Jazaeri, Farnaz, et al.
Veröffentlicht: (2025)
von: Jazaeri, Farnaz, et al.
Veröffentlicht: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
von: Close, George, et al.
Veröffentlicht: (2025)
von: Close, George, et al.
Veröffentlicht: (2025)
Speech Enhancement with Overlapped-Frame Information Fusion and Causal Self-Attention
von: Zhang, Yuewei, et al.
Veröffentlicht: (2025)
von: Zhang, Yuewei, et al.
Veröffentlicht: (2025)
A Multi-stage Low-latency Enhancement System for Hearing Aids
von: Ouyang, Chengwei, et al.
Veröffentlicht: (2025)
von: Ouyang, Chengwei, et al.
Veröffentlicht: (2025)
Harmonic Detection from Noisy Speech with Auditory Frame Gain for Intelligibility Enhancement
von: Queiroz, A., et al.
Veröffentlicht: (2024)
von: Queiroz, A., et al.
Veröffentlicht: (2024)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
von: Wisnu, Dyah A. M. G., et al.
Veröffentlicht: (2025)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
GESI: Gammachirp Envelope Similarity Index for Predicting Intelligibility of Simulated Hearing Loss Sounds
von: Yamamoto, Ayako, et al.
Veröffentlicht: (2023)
von: Yamamoto, Ayako, et al.
Veröffentlicht: (2023)
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
EventTrojan: Manipulating Non-Intrusive Speech Quality Assessment via Imperceptible Events
von: Ren, Ying, et al.
Veröffentlicht: (2023)
von: Ren, Ying, et al.
Veröffentlicht: (2023)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
von: Jo, Daejin, et al.
Veröffentlicht: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2021)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2021)
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2026)
Advances in Intelligent Hearing Aids: Deep Learning Approaches to Selective Noise Cancellation
von: Khan, Haris, et al.
Veröffentlicht: (2025)
von: Khan, Haris, et al.
Veröffentlicht: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024) -
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025) -
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024) -
Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
von: Zezario, Ryandhimas E.
Veröffentlicht: (2025) -
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)