Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ahn, Hoseong, Chae, Jeongyun, Park, Yoonji, Shim, Kyuhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
von: Park, Hansol, et al.
Veröffentlicht: (2025)
von: Park, Hansol, et al.
Veröffentlicht: (2025)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
von: Akinrintoyo, Emmanuel, et al.
Veröffentlicht: (2025)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2023)
von: Lin, Zhaofeng, et al.
Veröffentlicht: (2023)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
von: Liu, Wei, et al.
Veröffentlicht: (2023)
von: Liu, Wei, et al.
Veröffentlicht: (2023)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
von: Boccato, Tommaso, et al.
Veröffentlicht: (2026)
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
von: Honda, Tomoki, et al.
Veröffentlicht: (2024)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
Prompting Whisper for Joint Speech Transcription and Diarization
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
von: Zamyrova, Mariia, et al.
Veröffentlicht: (2026)
Probing Whisper for Dysarthric Speech in Detection and Assessment
von: Yue, Zhengjun, et al.
Veröffentlicht: (2025)
von: Yue, Zhengjun, et al.
Veröffentlicht: (2025)
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
von: Mu, Bingshen, et al.
Veröffentlicht: (2026)
von: Mu, Bingshen, et al.
Veröffentlicht: (2026)
FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
von: Lee, Junseok, et al.
Veröffentlicht: (2026)
von: Lee, Junseok, et al.
Veröffentlicht: (2026)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
A Study on Incorporating Whisper for Robust Speech Assessment
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
Mamba-based Decoder-Only Approach with Bidirectional Speech Modeling for Speech Recognition
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2024)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
von: Park, ChaeHun, et al.
Veröffentlicht: (2024)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
von: Fang, Zihao, et al.
Veröffentlicht: (2026)
von: Fang, Zihao, et al.
Veröffentlicht: (2026)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
von: Jogi, Yash, et al.
Veröffentlicht: (2025)
von: Jogi, Yash, et al.
Veröffentlicht: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
von: Yang, Hongli, et al.
Veröffentlicht: (2025)
von: Yang, Hongli, et al.
Veröffentlicht: (2025)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
von: Segal-Feldman, Yael, et al.
Veröffentlicht: (2024)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
von: Park, Hansol, et al.
Veröffentlicht: (2025) -
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026) -
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025) -
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
von: Wang, Shiyao, et al.
Veröffentlicht: (2025) -
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2025)