Listen Like a Teacher: Mitigating Whisper Hallucinations using Adaptive Layer Attention and Knowledge Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tripathi, Kumud, Menon, Aditya Srinivas, Gaurav, Aman, Gohil, Raj Prakash, Wasnik, Pankaj |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026)
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
von: Tripathi, Kumud, et al.
Veröffentlicht: (2025)
von: Tripathi, Kumud, et al.
Veröffentlicht: (2025)
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
von: Tripathi, Kumud, et al.
Veröffentlicht: (2024)
von: Tripathi, Kumud, et al.
Veröffentlicht: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Adaptive Knowledge Distillation using a Device-Aware Teacher for Low-Complexity Acoustic Scene Classification
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
von: Jeong, Seung Gyu, et al.
Veröffentlicht: (2025)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
von: Raj, Desh
Veröffentlicht: (2024)
von: Raj, Desh
Veröffentlicht: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
von: Lee, Junseok, et al.
Veröffentlicht: (2026)
von: Lee, Junseok, et al.
Veröffentlicht: (2026)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
von: Sahipjohn, Neha, et al.
Veröffentlicht: (2024)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
von: Sy, Yaya, et al.
Veröffentlicht: (2025)
von: Sy, Yaya, et al.
Veröffentlicht: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
von: Yuan, Xihao, et al.
Veröffentlicht: (2025)
von: Yuan, Xihao, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations
von: Bartolini, Jonatan, et al.
Veröffentlicht: (2024)
von: Bartolini, Jonatan, et al.
Veröffentlicht: (2024)
Noise-Aware In-Context Learning for Hallucination Mitigation in ALLMs
von: Huang, Qixuan, et al.
Veröffentlicht: (2026)
von: Huang, Qixuan, et al.
Veröffentlicht: (2026)
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
von: Pang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Pang, Jiacheng, et al.
Veröffentlicht: (2026)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2024)
von: Zhang, Li, et al.
Veröffentlicht: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Efficient Solutions for Mitigating Initialization Bias in Unsupervised Self-Adaptive Auditory Attention Decoding
von: Yao, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuanyuan, et al.
Veröffentlicht: (2025)
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
A Semi-Supervised Framework for Speech Confidence Detection using Whisper
von: Wynn, Adam, et al.
Veröffentlicht: (2026)
von: Wynn, Adam, et al.
Veröffentlicht: (2026)
DistilMOS: Layer-Wise Self-Distillation For Self-Supervised Learning Model-Based MOS Prediction
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering
von: Glazer, Neta, et al.
Veröffentlicht: (2026)
von: Glazer, Neta, et al.
Veröffentlicht: (2026)
Adaptive Knowledge Distillation for Device-Directed Speech Detection
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
von: Agrawal, Saurabh, et al.
Veröffentlicht: (2025)
von: Agrawal, Saurabh, et al.
Veröffentlicht: (2025)
Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
von: Ioannides, Georgios, et al.
Veröffentlicht: (2024)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2024)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
von: Bijoy, Mehedi Hasan, et al.
Veröffentlicht: (2025)
von: Bijoy, Mehedi Hasan, et al.
Veröffentlicht: (2025)
Target Speaker ASR with Whisper
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
von: Shen, Yu-Han, et al.
Veröffentlicht: (2018)
von: Shen, Yu-Han, et al.
Veröffentlicht: (2018)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025) -
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2026) -
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
von: Tripathi, Kumud, et al.
Veröffentlicht: (2025) -
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
von: Tripathi, Kumud, et al.
Veröffentlicht: (2024) -
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)