A Semi-Supervised Framework for Speech Confidence Detection using Whisper
Fuente:
arXiv
Saved in:
| Main Authors: | Wynn, Adam, Wang, Jingyun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
by: Adnan, Tariq, et al.
Published: (2024)
by: Adnan, Tariq, et al.
Published: (2024)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
by: Zezario, Ryandhimas E., et al.
Published: (2023)
by: Zezario, Ryandhimas E., et al.
Published: (2023)
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
by: Chowdhury, Aurchi, et al.
Published: (2026)
by: Chowdhury, Aurchi, et al.
Published: (2026)
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
by: Justus, Aju Ani, et al.
Published: (2026)
by: Justus, Aju Ani, et al.
Published: (2026)
Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition Systems
by: Jin, Weifei, et al.
Published: (2025)
by: Jin, Weifei, et al.
Published: (2025)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
by: Ma, Yujian, et al.
Published: (2025)
by: Ma, Yujian, et al.
Published: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
by: Close, George, et al.
Published: (2025)
by: Close, George, et al.
Published: (2025)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
by: Wang, Charles L., et al.
Published: (2026)
by: Wang, Charles L., et al.
Published: (2026)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
by: Ameer, Huma, et al.
Published: (2024)
by: Ameer, Huma, et al.
Published: (2024)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
by: Shendabadi, Ali, et al.
Published: (2026)
by: Shendabadi, Ali, et al.
Published: (2026)
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
by: Ravenscroft, William, et al.
Published: (2025)
by: Ravenscroft, William, et al.
Published: (2025)
Assessing the Impact of Speaker Identity in Speech Spoofing Detection
by: Dao, Anh-Tuan, et al.
Published: (2026)
by: Dao, Anh-Tuan, et al.
Published: (2026)
PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
by: Pahar, Madhurananda, et al.
Published: (2026)
by: Pahar, Madhurananda, et al.
Published: (2026)
WhisperRT -- Turning Whisper into a Causal Streaming Model
by: Krichli, Tomer, et al.
Published: (2025)
by: Krichli, Tomer, et al.
Published: (2025)
Investigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments
by: Anacin, et al.
Published: (2026)
by: Anacin, et al.
Published: (2026)
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
by: Goswami, Mandip
Published: (2026)
by: Goswami, Mandip
Published: (2026)
Semantic-Aware Confidence Calibration for Automated Audio Captioning
by: Dunker, Lucas, et al.
Published: (2025)
by: Dunker, Lucas, et al.
Published: (2025)
Can DeepFake Speech be Reliably Detected?
by: Liu, Hongbin, et al.
Published: (2024)
by: Liu, Hongbin, et al.
Published: (2024)
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
by: Sohn, Samuel S., et al.
Published: (2025)
by: Sohn, Samuel S., et al.
Published: (2025)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
by: Kumar, Gokul Karthik, et al.
Published: (2026)
by: Kumar, Gokul Karthik, et al.
Published: (2026)
When Denoising Hinders: Revisiting Zero-Shot ASR with SAM-Audio and Whisper
by: Islam, Akif, et al.
Published: (2026)
by: Islam, Akif, et al.
Published: (2026)
Multi-Channel Replay Speech Detection using Acoustic Maps
by: Neri, Michael, et al.
Published: (2026)
by: Neri, Michael, et al.
Published: (2026)
Enhancing Lung Disease Diagnosis via Semi-Supervised Machine Learning
by: Xu, Xiaoran, et al.
Published: (2025)
by: Xu, Xiaoran, et al.
Published: (2025)
Whispy: Adapting STT Whisper Models to Real-Time Environments
by: Bevilacqua, Antonio, et al.
Published: (2024)
by: Bevilacqua, Antonio, et al.
Published: (2024)
Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering
by: Tran, Dinh Phu, et al.
Published: (2026)
by: Tran, Dinh Phu, et al.
Published: (2026)
A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning
by: Ilias, Loukas, et al.
Published: (2026)
by: Ilias, Loukas, et al.
Published: (2026)
Impact of Speech Mode in Automatic Pathological Speech Detection
by: Sheikh, Shakeel A., et al.
Published: (2024)
by: Sheikh, Shakeel A., et al.
Published: (2024)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
by: Fu, Szu-Wei, et al.
Published: (2024)
by: Fu, Szu-Wei, et al.
Published: (2024)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
by: Fu, Yonggan, et al.
Published: (2022)
by: Fu, Yonggan, et al.
Published: (2022)
Music Transcription with (Almost) No Supervision
by: Shin, Saebyeol, et al.
Published: (2026)
by: Shin, Saebyeol, et al.
Published: (2026)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
by: Vaessen, Nik, et al.
Published: (2024)
by: Vaessen, Nik, et al.
Published: (2024)
Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias
by: Fursule, Aishwarya, et al.
Published: (2026)
by: Fursule, Aishwarya, et al.
Published: (2026)
Masked Autoencoders as Universal Speech Enhancer
by: Rajagopalan, Rajalaxmi, et al.
Published: (2026)
by: Rajagopalan, Rajalaxmi, et al.
Published: (2026)
Detecting Throat Cancer from Speech Signals using Machine Learning: A Scoping Literature Review
by: Paterson, Mary, et al.
Published: (2023)
by: Paterson, Mary, et al.
Published: (2023)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
by: Brima, Yusuf, et al.
Published: (2023)
by: Brima, Yusuf, et al.
Published: (2023)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
by: Feng, Pengchao, et al.
Published: (2025)
by: Feng, Pengchao, et al.
Published: (2025)
A Framework for Evaluating Faithfulness in Explainable AI for Machine Anomalous Sound Detection Using Frequency-Band Perturbation
by: Buck, Alexander, et al.
Published: (2026)
by: Buck, Alexander, et al.
Published: (2026)
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
by: Kumar, Sahil, et al.
Published: (2026)
by: Kumar, Sahil, et al.
Published: (2026)
Similar Items
-
A Novel Fusion Architecture for PD Detection Using Semi-Supervised Speech Embeddings
by: Adnan, Tariq, et al.
Published: (2024) -
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025) -
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
by: Zezario, Ryandhimas E., et al.
Published: (2023) -
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
by: Chowdhury, Aurchi, et al.
Published: (2026) -
voice2mode: Phonation Mode Classification in Singing using Self-Supervised Speech Models
by: Justus, Aju Ani, et al.
Published: (2026)