Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
Fuente:
arXiv
Salvato in:
| Autori principali: | Segal-Feldman, Yael, Shamsian, Aviv, Navon, Aviv, Hetz, Gill, Keshet, Joseph |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlowTSE: Target Speaker Extraction with Flow Matching
di: Navon, Aviv, et al.
Pubblicazione: (2025)
di: Navon, Aviv, et al.
Pubblicazione: (2025)
Drax: Speech Recognition with Discrete Flow Matching
di: Navon, Aviv, et al.
Pubblicazione: (2025)
di: Navon, Aviv, et al.
Pubblicazione: (2025)
Keyword-Guided Adaptation of Automatic Speech Recognition
di: Shamsian, Aviv, et al.
Pubblicazione: (2024)
di: Shamsian, Aviv, et al.
Pubblicazione: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2025)
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2025)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
di: Yemini, Yochai, et al.
Pubblicazione: (2023)
di: Yemini, Yochai, et al.
Pubblicazione: (2023)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
di: Guo, Pengcheng, et al.
Pubblicazione: (2024)
di: Guo, Pengcheng, et al.
Pubblicazione: (2024)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
di: Boccato, Tommaso, et al.
Pubblicazione: (2026)
di: Boccato, Tommaso, et al.
Pubblicazione: (2026)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
di: Barański, Mateusz, et al.
Pubblicazione: (2025)
di: Barański, Mateusz, et al.
Pubblicazione: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
di: Ahn, Hoseong, et al.
Pubblicazione: (2026)
di: Ahn, Hoseong, et al.
Pubblicazione: (2026)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
di: Orhon, Atila, et al.
Pubblicazione: (2025)
di: Orhon, Atila, et al.
Pubblicazione: (2025)
Efficient Scaling for LLM-based ASR
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
di: Zhao, Wenbo, et al.
Pubblicazione: (2024)
di: Zhao, Wenbo, et al.
Pubblicazione: (2024)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
di: Song, Zheshu, et al.
Pubblicazione: (2024)
di: Song, Zheshu, et al.
Pubblicazione: (2024)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
di: Farhadipour, Aref, et al.
Pubblicazione: (2025)
di: Farhadipour, Aref, et al.
Pubblicazione: (2025)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
di: Kwon, Taeyoun, et al.
Pubblicazione: (2025)
di: Kwon, Taeyoun, et al.
Pubblicazione: (2025)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
di: Fang, Zihao, et al.
Pubblicazione: (2026)
di: Fang, Zihao, et al.
Pubblicazione: (2026)
MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding
di: Xi, Yu, et al.
Pubblicazione: (2025)
di: Xi, Yu, et al.
Pubblicazione: (2025)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
di: Wei, Linye, et al.
Pubblicazione: (2025)
di: Wei, Linye, et al.
Pubblicazione: (2025)
Extending Whisper with prompt tuning to target-speaker ASR
di: Ma, Hao, et al.
Pubblicazione: (2023)
di: Ma, Hao, et al.
Pubblicazione: (2023)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
di: Zhao, Yiyang, et al.
Pubblicazione: (2024)
di: Zhao, Yiyang, et al.
Pubblicazione: (2024)
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
di: Turetzky, Arnon, et al.
Pubblicazione: (2024)
di: Turetzky, Arnon, et al.
Pubblicazione: (2024)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
di: Wang, Weiqing, et al.
Pubblicazione: (2024)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
di: Bhuiyan, Mohammed Aman, et al.
Pubblicazione: (2026)
di: Bhuiyan, Mohammed Aman, et al.
Pubblicazione: (2026)
Unifying Streaming and Non-streaming Zipformer-based ASR
di: Sharma, Bidisha, et al.
Pubblicazione: (2025)
di: Sharma, Bidisha, et al.
Pubblicazione: (2025)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
di: Chen, Qian, et al.
Pubblicazione: (2023)
di: Chen, Qian, et al.
Pubblicazione: (2023)
Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
di: Ho, Tuan Vu, et al.
Pubblicazione: (2025)
di: Ho, Tuan Vu, et al.
Pubblicazione: (2025)
Deepfake Detection of Singing Voices With Whisper Encodings
di: Sharma, Falguni, et al.
Pubblicazione: (2025)
di: Sharma, Falguni, et al.
Pubblicazione: (2025)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
di: Xu, Tianyi, et al.
Pubblicazione: (2024)
di: Xu, Tianyi, et al.
Pubblicazione: (2024)
Whisper-RIR-Mega: A Paired Clean-Reverberant Speech Benchmark for ASR Robustness to Room Acoustics
di: Goswami, Mandip
Pubblicazione: (2026)
di: Goswami, Mandip
Pubblicazione: (2026)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
di: Dissen, Yehoshua, et al.
Pubblicazione: (2024)
di: Dissen, Yehoshua, et al.
Pubblicazione: (2024)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
di: Zhang, Li, et al.
Pubblicazione: (2024)
di: Zhang, Li, et al.
Pubblicazione: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FlowTSE: Target Speaker Extraction with Flow Matching
di: Navon, Aviv, et al.
Pubblicazione: (2025) -
Drax: Speech Recognition with Discrete Flow Matching
di: Navon, Aviv, et al.
Pubblicazione: (2025) -
Keyword-Guided Adaptation of Automatic Speech Recognition
di: Shamsian, Aviv, et al.
Pubblicazione: (2024) -
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025) -
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
di: Glazer, Neta, et al.
Pubblicazione: (2025)