Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haoyu, Hu, Guoqiang, Lin, Guodong, Zhang, Wei-Qiang, Li, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Extending Whisper with prompt tuning to target-speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2023)
von: Ma, Hao, et al.
Veröffentlicht: (2023)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
von: Sy, Yaya, et al.
Veröffentlicht: (2025)
von: Sy, Yaya, et al.
Veröffentlicht: (2025)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
Transfer Learning from Whisper for Microscopic Intelligibility Prediction
von: Best, Paul, et al.
Veröffentlicht: (2024)
von: Best, Paul, et al.
Veröffentlicht: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
von: Zhuo, Le, et al.
Veröffentlicht: (2023)
kNN For Whisper And Its Effect On Bias And Speaker Adaptation
von: Nachesa, Maya K., et al.
Veröffentlicht: (2024)
von: Nachesa, Maya K., et al.
Veröffentlicht: (2024)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2024)
von: Zhang, Li, et al.
Veröffentlicht: (2024)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2024)
Quantizing Whisper-small: How design choices affect ASR performance
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
von: Söhler, Arthur, et al.
Veröffentlicht: (2025)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
von: Orhon, Atila, et al.
Veröffentlicht: (2025)
von: Orhon, Atila, et al.
Veröffentlicht: (2025)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
von: Özyilmaz, Ömer Tarik, et al.
Veröffentlicht: (2025)
von: Özyilmaz, Ömer Tarik, et al.
Veröffentlicht: (2025)
Audio-to-Score Conversion Model Based on Whisper methodology
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhao, Yiyang, et al.
Veröffentlicht: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
Probing Whisper for Dysarthric Speech in Detection and Assessment
von: Yue, Zhengjun, et al.
Veröffentlicht: (2025)
von: Yue, Zhengjun, et al.
Veröffentlicht: (2025)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025) -
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024) -
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024) -
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025) -
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)