WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
Fuente:
arXiv
Salvato in:
| Autori principali: | Close, George, Hong, Kris, Hain, Thomas, Goetze, Stefan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
di: Close, George, et al.
Pubblicazione: (2024)
di: Close, George, et al.
Pubblicazione: (2024)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024)
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
di: Ravenscroft, William, et al.
Pubblicazione: (2024)
di: Ravenscroft, William, et al.
Pubblicazione: (2024)
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
di: Ravenscroft, William, et al.
Pubblicazione: (2025)
di: Ravenscroft, William, et al.
Pubblicazione: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
di: Yamamoto, Katsuhiko, et al.
Pubblicazione: (2025)
di: Yamamoto, Katsuhiko, et al.
Pubblicazione: (2025)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
di: Ameer, Huma, et al.
Pubblicazione: (2024)
di: Ameer, Huma, et al.
Pubblicazione: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
di: Nakazawa, Kazushi
Pubblicazione: (2026)
di: Nakazawa, Kazushi
Pubblicazione: (2026)
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Non-Intrusive Intelligibility Prediction for Hearing Aids: Recent Advances, Trends, and Challenges
di: Zezario, Ryandhimas E.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E.
Pubblicazione: (2025)
LIWhiz: A Non-Intrusive Lyric Intelligibility Prediction System for the Cadenza Challenge
di: Shekar, Ram C. M. C., et al.
Pubblicazione: (2025)
di: Shekar, Ram C. M. C., et al.
Pubblicazione: (2025)
EventTrojan: Manipulating Non-Intrusive Speech Quality Assessment via Imperceptible Events
di: Ren, Ying, et al.
Pubblicazione: (2023)
di: Ren, Ying, et al.
Pubblicazione: (2023)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
di: Zhao, Yiyang, et al.
Pubblicazione: (2024)
di: Zhao, Yiyang, et al.
Pubblicazione: (2024)
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
di: Cao, Boxuan, et al.
Pubblicazione: (2025)
di: Cao, Boxuan, et al.
Pubblicazione: (2025)
Prompting Whisper for Joint Speech Transcription and Diarization
di: Zamyrova, Mariia, et al.
Pubblicazione: (2026)
di: Zamyrova, Mariia, et al.
Pubblicazione: (2026)
Probing Whisper for Dysarthric Speech in Detection and Assessment
di: Yue, Zhengjun, et al.
Pubblicazione: (2025)
di: Yue, Zhengjun, et al.
Pubblicazione: (2025)
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
di: Wagner, Dominik, et al.
Pubblicazione: (2023)
di: Wagner, Dominik, et al.
Pubblicazione: (2023)
Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech
di: de Oliveira, Danilo, et al.
Pubblicazione: (2024)
di: de Oliveira, Danilo, et al.
Pubblicazione: (2024)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
di: Ma, Yujian, et al.
Pubblicazione: (2025)
di: Ma, Yujian, et al.
Pubblicazione: (2025)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
di: Farhadipour, Aref, et al.
Pubblicazione: (2024)
di: Farhadipour, Aref, et al.
Pubblicazione: (2024)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
di: Fang, Zihao, et al.
Pubblicazione: (2026)
di: Fang, Zihao, et al.
Pubblicazione: (2026)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
di: Akinrintoyo, Emmanuel, et al.
Pubblicazione: (2025)
di: Akinrintoyo, Emmanuel, et al.
Pubblicazione: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
di: Park, Chanho, et al.
Pubblicazione: (2023)
di: Park, Chanho, et al.
Pubblicazione: (2023)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
di: Farhadipour, Aref, et al.
Pubblicazione: (2025)
di: Farhadipour, Aref, et al.
Pubblicazione: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
di: Clarke, Jason, et al.
Pubblicazione: (2025)
di: Clarke, Jason, et al.
Pubblicazione: (2025)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
di: Boccato, Tommaso, et al.
Pubblicazione: (2026)
di: Boccato, Tommaso, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
di: Close, George, et al.
Pubblicazione: (2024) -
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024) -
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
di: Sutherland, Robert, et al.
Pubblicazione: (2024) -
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
di: Ravenscroft, William, et al.
Pubblicazione: (2024) -
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
di: Ravenscroft, William, et al.
Pubblicazione: (2025)