Adopting Whisper for Confidence Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Aggarwal, Vaibhav, Nair, Shabari S, Verma, Yash, Jogi, Yash |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
by: Jogi, Yash, et al.
Published: (2025)
by: Jogi, Yash, et al.
Published: (2025)
MaskCycleGAN-based Whisper to Normal Speech Conversion
by: Gupta, K. Rohith, et al.
Published: (2024)
by: Gupta, K. Rohith, et al.
Published: (2024)
RiTTA: Modeling Event Relations in Text-to-Audio Generation
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
by: Polok, Alexander, et al.
Published: (2026)
by: Polok, Alexander, et al.
Published: (2026)
Investigating Confidence Estimation Measures for Speaker Diarization
by: Chowdhury, Anurag, et al.
Published: (2024)
by: Chowdhury, Anurag, et al.
Published: (2024)
Quartered Spectral Envelope and 1D-CNN-based Classification of Normally Phonated and Whispered Speech
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
by: Sohn, Samuel S., et al.
Published: (2025)
by: Sohn, Samuel S., et al.
Published: (2025)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
by: Verma, Prateek
Published: (2024)
by: Verma, Prateek
Published: (2024)
Whispy: Adapting STT Whisper Models to Real-Time Environments
by: Bevilacqua, Antonio, et al.
Published: (2024)
by: Bevilacqua, Antonio, et al.
Published: (2024)
Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
by: Ravi, Nagarathna, et al.
Published: (2024)
by: Ravi, Nagarathna, et al.
Published: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
by: Krichli, Tomer, et al.
Published: (2025)
by: Krichli, Tomer, et al.
Published: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
by: Zezario, Ryandhimas E., et al.
Published: (2023)
by: Zezario, Ryandhimas E., et al.
Published: (2023)
Expressive Timing in Hindustani Vocal Music
by: Bhake, Yash, et al.
Published: (2025)
by: Bhake, Yash, et al.
Published: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
by: Ma, Yujian, et al.
Published: (2025)
by: Ma, Yujian, et al.
Published: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
by: Close, George, et al.
Published: (2025)
by: Close, George, et al.
Published: (2025)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
by: Ameer, Huma, et al.
Published: (2024)
by: Ameer, Huma, et al.
Published: (2024)
A Language Model With Million Context Length For Raw Audio
by: Verma, Prateek
Published: (2022)
by: Verma, Prateek
Published: (2022)
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
by: Ravenscroft, William, et al.
Published: (2025)
by: Ravenscroft, William, et al.
Published: (2025)
Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
by: Wong, Zheng Jie, et al.
Published: (2025)
by: Wong, Zheng Jie, et al.
Published: (2025)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Audio-to-Score Conversion Model Based on Whisper methodology
by: Zhang, Hongyao, et al.
Published: (2024)
by: Zhang, Hongyao, et al.
Published: (2024)
TellWhisper: Tell Whisper Who Speaks When
by: Hu, Yifan, et al.
Published: (2026)
by: Hu, Yifan, et al.
Published: (2026)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
by: Segal-Feldman, Yael, et al.
Published: (2024)
by: Segal-Feldman, Yael, et al.
Published: (2024)
Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier
by: Hayta, Berk, et al.
Published: (2026)
by: Hayta, Berk, et al.
Published: (2026)
Melodic and Metrical Elements of Expressiveness in Hindustani Vocal Music
by: Bhake, Yash, et al.
Published: (2025)
by: Bhake, Yash, et al.
Published: (2025)
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
by: Goswami, Mandip
Published: (2026)
by: Goswami, Mandip
Published: (2026)
Acoustics-specific Piano Velocity Estimation
by: Simonetta, Federico, et al.
Published: (2022)
by: Simonetta, Federico, et al.
Published: (2022)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
by: Wu, Wenxuan, et al.
Published: (2025)
by: Wu, Wenxuan, et al.
Published: (2025)
Sound Event Detection and Localization with Distance Estimation
by: Krause, Daniel Aleksander, et al.
Published: (2024)
by: Krause, Daniel Aleksander, et al.
Published: (2024)
Toward Fully Self-Supervised Multi-Pitch Estimation
by: Cwitkowitz, Frank, et al.
Published: (2024)
by: Cwitkowitz, Frank, et al.
Published: (2024)
HRTF Estimation using a Score-based Prior
by: Thuillier, Etienne, et al.
Published: (2024)
by: Thuillier, Etienne, et al.
Published: (2024)
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
Enhancing Aviation Communication Transcription: Fine-Tuning Distil-Whisper with LoRA
by: Mirzaei, Shokoufeh, et al.
Published: (2025)
by: Mirzaei, Shokoufeh, et al.
Published: (2025)
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
by: Cwitkowitz, Frank, et al.
Published: (2025)
by: Cwitkowitz, Frank, et al.
Published: (2025)
Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation
by: Nie, Jingping, et al.
Published: (2025)
by: Nie, Jingping, et al.
Published: (2025)
Maximum Likelihood Estimation of the Direction of Sound In A Reverberant Noisy Environment
by: Mansour, Mohamed F.
Published: (2024)
by: Mansour, Mohamed F.
Published: (2024)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
by: Primus, Paul, et al.
Published: (2024)
by: Primus, Paul, et al.
Published: (2024)
Similar Items
-
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
by: Jogi, Yash, et al.
Published: (2025) -
MaskCycleGAN-based Whisper to Normal Speech Conversion
by: Gupta, K. Rohith, et al.
Published: (2024) -
RiTTA: Modeling Event Relations in Text-to-Audio Generation
by: He, Yuhang, et al.
Published: (2024) -
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
by: Polok, Alexander, et al.
Published: (2026) -
Investigating Confidence Estimation Measures for Speaker Diarization
by: Chowdhury, Anurag, et al.
Published: (2024)