Adopting Whisper for Confidence Estimation
Fuente:
arXiv
Guardado en:
| Autores principales: | Aggarwal, Vaibhav, Nair, Shabari S, Verma, Yash, Jogi, Yash |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
por: Jogi, Yash, et al.
Publicado: (2025)
por: Jogi, Yash, et al.
Publicado: (2025)
MaskCycleGAN-based Whisper to Normal Speech Conversion
por: Gupta, K. Rohith, et al.
Publicado: (2024)
por: Gupta, K. Rohith, et al.
Publicado: (2024)
RiTTA: Modeling Event Relations in Text-to-Audio Generation
por: He, Yuhang, et al.
Publicado: (2024)
por: He, Yuhang, et al.
Publicado: (2024)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
por: Polok, Alexander, et al.
Publicado: (2026)
por: Polok, Alexander, et al.
Publicado: (2026)
Investigating Confidence Estimation Measures for Speaker Diarization
por: Chowdhury, Anurag, et al.
Publicado: (2024)
por: Chowdhury, Anurag, et al.
Publicado: (2024)
Quartered Spectral Envelope and 1D-CNN-based Classification of Normally Phonated and Whispered Speech
por: Joysingh, S. Johanan, et al.
Publicado: (2024)
por: Joysingh, S. Johanan, et al.
Publicado: (2024)
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
por: Sohn, Samuel S., et al.
Publicado: (2025)
por: Sohn, Samuel S., et al.
Publicado: (2025)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
por: Verma, Prateek
Publicado: (2024)
por: Verma, Prateek
Publicado: (2024)
Whispy: Adapting STT Whisper Models to Real-Time Environments
por: Bevilacqua, Antonio, et al.
Publicado: (2024)
por: Bevilacqua, Antonio, et al.
Publicado: (2024)
Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification
por: Joysingh, S. Johanan, et al.
Publicado: (2024)
por: Joysingh, S. Johanan, et al.
Publicado: (2024)
TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
por: Ravi, Nagarathna, et al.
Publicado: (2024)
por: Ravi, Nagarathna, et al.
Publicado: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
Expressive Timing in Hindustani Vocal Music
por: Bhake, Yash, et al.
Publicado: (2025)
por: Bhake, Yash, et al.
Publicado: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
por: Ma, Yujian, et al.
Publicado: (2025)
por: Ma, Yujian, et al.
Publicado: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
por: Close, George, et al.
Publicado: (2025)
por: Close, George, et al.
Publicado: (2025)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
por: Ameer, Huma, et al.
Publicado: (2024)
por: Ameer, Huma, et al.
Publicado: (2024)
A Language Model With Million Context Length For Raw Audio
por: Verma, Prateek
Publicado: (2022)
por: Verma, Prateek
Publicado: (2022)
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
por: Ravenscroft, William, et al.
Publicado: (2025)
por: Ravenscroft, William, et al.
Publicado: (2025)
Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
por: Wong, Zheng Jie, et al.
Publicado: (2025)
por: Wong, Zheng Jie, et al.
Publicado: (2025)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
por: Mancini, Eleonora, et al.
Publicado: (2025)
por: Mancini, Eleonora, et al.
Publicado: (2025)
Audio-to-Score Conversion Model Based on Whisper methodology
por: Zhang, Hongyao, et al.
Publicado: (2024)
por: Zhang, Hongyao, et al.
Publicado: (2024)
TellWhisper: Tell Whisper Who Speaks When
por: Hu, Yifan, et al.
Publicado: (2026)
por: Hu, Yifan, et al.
Publicado: (2026)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
por: Segal-Feldman, Yael, et al.
Publicado: (2024)
por: Segal-Feldman, Yael, et al.
Publicado: (2024)
Plug-in Losses for Evidential Deep Learning: A Simplified Framework for Uncertainty Estimation that Includes the Softmax Classifier
por: Hayta, Berk, et al.
Publicado: (2026)
por: Hayta, Berk, et al.
Publicado: (2026)
Melodic and Metrical Elements of Expressiveness in Hindustani Vocal Music
por: Bhake, Yash, et al.
Publicado: (2025)
por: Bhake, Yash, et al.
Publicado: (2025)
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
por: Goswami, Mandip
Publicado: (2026)
por: Goswami, Mandip
Publicado: (2026)
Acoustics-specific Piano Velocity Estimation
por: Simonetta, Federico, et al.
Publicado: (2022)
por: Simonetta, Federico, et al.
Publicado: (2022)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
por: Wu, Wenxuan, et al.
Publicado: (2025)
por: Wu, Wenxuan, et al.
Publicado: (2025)
Sound Event Detection and Localization with Distance Estimation
por: Krause, Daniel Aleksander, et al.
Publicado: (2024)
por: Krause, Daniel Aleksander, et al.
Publicado: (2024)
Toward Fully Self-Supervised Multi-Pitch Estimation
por: Cwitkowitz, Frank, et al.
Publicado: (2024)
por: Cwitkowitz, Frank, et al.
Publicado: (2024)
HRTF Estimation using a Score-based Prior
por: Thuillier, Etienne, et al.
Publicado: (2024)
por: Thuillier, Etienne, et al.
Publicado: (2024)
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
por: Nareddy, Kartheek Kumar Reddy, et al.
Publicado: (2025)
por: Nareddy, Kartheek Kumar Reddy, et al.
Publicado: (2025)
Enhancing Aviation Communication Transcription: Fine-Tuning Distil-Whisper with LoRA
por: Mirzaei, Shokoufeh, et al.
Publicado: (2025)
por: Mirzaei, Shokoufeh, et al.
Publicado: (2025)
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
por: Cwitkowitz, Frank, et al.
Publicado: (2025)
por: Cwitkowitz, Frank, et al.
Publicado: (2025)
Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation
por: Nie, Jingping, et al.
Publicado: (2025)
por: Nie, Jingping, et al.
Publicado: (2025)
Maximum Likelihood Estimation of the Direction of Sound In A Reverberant Noisy Environment
por: Mansour, Mohamed F.
Publicado: (2024)
por: Mansour, Mohamed F.
Publicado: (2024)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
por: Primus, Paul, et al.
Publicado: (2024)
por: Primus, Paul, et al.
Publicado: (2024)
Ejemplares similares
-
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
por: Jogi, Yash, et al.
Publicado: (2025) -
MaskCycleGAN-based Whisper to Normal Speech Conversion
por: Gupta, K. Rohith, et al.
Publicado: (2024) -
RiTTA: Modeling Event Relations in Text-to-Audio Generation
por: He, Yuhang, et al.
Publicado: (2024) -
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
por: Polok, Alexander, et al.
Publicado: (2026) -
Investigating Confidence Estimation Measures for Speaker Diarization
por: Chowdhury, Anurag, et al.
Publicado: (2024)