One Whisper to Grade Them All
Fuente:
arXiv
Salvato in:
| Autori principali: | Phan, Nhan, Porwal, Anusha, Getman, Yaroslav, Voskoboinik, Ekaterina, Grósz, Tamás, Kurimo, Mikko |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
di: Getman, Yaroslav, et al.
Pubblicazione: (2025)
di: Getman, Yaroslav, et al.
Pubblicazione: (2025)
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
di: Phan, Nhan, et al.
Pubblicazione: (2025)
di: Phan, Nhan, et al.
Pubblicazione: (2025)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025)
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025)
Out-of-distribution generalisation in spoken language understanding
di: Porjazovski, Dejan, et al.
Pubblicazione: (2024)
di: Porjazovski, Dejan, et al.
Pubblicazione: (2024)
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
di: Li, Zirui, et al.
Pubblicazione: (2025)
di: Li, Zirui, et al.
Pubblicazione: (2025)
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
di: Li, Zirui, et al.
Pubblicazione: (2025)
di: Li, Zirui, et al.
Pubblicazione: (2025)
Whisper Has an Internal Word Aligner
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2025)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2025)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
PhoWhisper: Automatic Speech Recognition for Vietnamese
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
di: Le, Thanh-Thien, et al.
Pubblicazione: (2024)
HydraFormer: One Encoder For All Subsampling Rates
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
Fine-tuning Whisper on Low-Resource Languages for Real-World Applications
di: Timmel, Vincenzo, et al.
Pubblicazione: (2024)
di: Timmel, Vincenzo, et al.
Pubblicazione: (2024)
Enhancing Whisper's Accuracy and Speed for Indian Languages through Prompt-Tuning and Tokenization
di: Tripathi, Kumud, et al.
Pubblicazione: (2024)
di: Tripathi, Kumud, et al.
Pubblicazione: (2024)
Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper
di: Tran, Hoan My, et al.
Pubblicazione: (2026)
di: Tran, Hoan My, et al.
Pubblicazione: (2026)
Transfer Learning from Whisper for Microscopic Intelligibility Prediction
di: Best, Paul, et al.
Pubblicazione: (2024)
di: Best, Paul, et al.
Pubblicazione: (2024)
Can Whisper perform speech-based in-context learning?
di: Wang, Siyin, et al.
Pubblicazione: (2023)
di: Wang, Siyin, et al.
Pubblicazione: (2023)
Extending Whisper with prompt tuning to target-speaker ASR
di: Ma, Hao, et al.
Pubblicazione: (2023)
di: Ma, Hao, et al.
Pubblicazione: (2023)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
kNN For Whisper And Its Effect On Bias And Speaker Adaptation
di: Nachesa, Maya K., et al.
Pubblicazione: (2024)
di: Nachesa, Maya K., et al.
Pubblicazione: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
di: Hu, Rui, et al.
Pubblicazione: (2025)
di: Hu, Rui, et al.
Pubblicazione: (2025)
Quantizing Whisper-small: How design choices affect ASR performance
di: Söhler, Arthur, et al.
Pubblicazione: (2025)
di: Söhler, Arthur, et al.
Pubblicazione: (2025)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
di: Orhon, Atila, et al.
Pubblicazione: (2025)
di: Orhon, Atila, et al.
Pubblicazione: (2025)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
di: Shao, Hang, et al.
Pubblicazione: (2023)
di: Shao, Hang, et al.
Pubblicazione: (2023)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
di: Yang, Chih-Kai, et al.
Pubblicazione: (2024)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
di: Sy, Yaya, et al.
Pubblicazione: (2025)
di: Sy, Yaya, et al.
Pubblicazione: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
di: Vesterbacka, Leonora, et al.
Pubblicazione: (2025)
di: Vesterbacka, Leonora, et al.
Pubblicazione: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
di: Kocour, Martin, et al.
Pubblicazione: (2025)
di: Kocour, Martin, et al.
Pubblicazione: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
di: Zhao, Jiahui, et al.
Pubblicazione: (2024)
di: Zhao, Jiahui, et al.
Pubblicazione: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
di: Zhuo, Le, et al.
Pubblicazione: (2023)
di: Zhuo, Le, et al.
Pubblicazione: (2023)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
di: Getman, Yaroslav, et al.
Pubblicazione: (2025) -
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
di: Phan, Nhan, et al.
Pubblicazione: (2025) -
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025) -
Out-of-distribution generalisation in spoken language understanding
di: Porjazovski, Dejan, et al.
Pubblicazione: (2024) -
Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech
di: Li, Zirui, et al.
Pubblicazione: (2025)