Salvato in:
| Autori principali: | Best, Paul, Cuervo, Santiago, Marxer, Ricard |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2404.01737 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speech foundation models on intelligibility prediction for hearing-impaired listeners
di: Cuervo, Santiago, et al.
Pubblicazione: (2024)
di: Cuervo, Santiago, et al.
Pubblicazione: (2024)
Scaling Properties of Speech Language Models
di: Cuervo, Santiago, et al.
Pubblicazione: (2024)
di: Cuervo, Santiago, et al.
Pubblicazione: (2024)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
di: Cuervo, Santiago, et al.
Pubblicazione: (2021)
di: Cuervo, Santiago, et al.
Pubblicazione: (2021)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)
Can Whisper perform speech-based in-context learning?
di: Wang, Siyin, et al.
Pubblicazione: (2023)
di: Wang, Siyin, et al.
Pubblicazione: (2023)
Extending Whisper with prompt tuning to target-speaker ASR
di: Ma, Hao, et al.
Pubblicazione: (2023)
di: Ma, Hao, et al.
Pubblicazione: (2023)
kNN For Whisper And Its Effect On Bias And Speaker Adaptation
di: Nachesa, Maya K., et al.
Pubblicazione: (2024)
di: Nachesa, Maya K., et al.
Pubblicazione: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper
di: Yang, Chih-Kai, et al.
Pubblicazione: (2024)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
di: Hu, Rui, et al.
Pubblicazione: (2025)
di: Hu, Rui, et al.
Pubblicazione: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
di: Shao, Hang, et al.
Pubblicazione: (2023)
di: Shao, Hang, et al.
Pubblicazione: (2023)
Quantizing Whisper-small: How design choices affect ASR performance
di: Söhler, Arthur, et al.
Pubblicazione: (2025)
di: Söhler, Arthur, et al.
Pubblicazione: (2025)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
di: Orhon, Atila, et al.
Pubblicazione: (2025)
di: Orhon, Atila, et al.
Pubblicazione: (2025)
Factorized RVQ-GAN For Disentangled Speech Tokenization
di: Khurana, Sameer, et al.
Pubblicazione: (2025)
di: Khurana, Sameer, et al.
Pubblicazione: (2025)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
di: Zhao, Jiahui, et al.
Pubblicazione: (2024)
di: Zhao, Jiahui, et al.
Pubblicazione: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
di: Raina, Vyas, et al.
Pubblicazione: (2024)
di: Raina, Vyas, et al.
Pubblicazione: (2024)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
di: Vesterbacka, Leonora, et al.
Pubblicazione: (2025)
di: Vesterbacka, Leonora, et al.
Pubblicazione: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
di: Kocour, Martin, et al.
Pubblicazione: (2025)
di: Kocour, Martin, et al.
Pubblicazione: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
di: Sy, Yaya, et al.
Pubblicazione: (2025)
di: Sy, Yaya, et al.
Pubblicazione: (2025)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
di: Cui, Ziyun, et al.
Pubblicazione: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
di: Zhuo, Le, et al.
Pubblicazione: (2023)
di: Zhuo, Le, et al.
Pubblicazione: (2023)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
di: Özyilmaz, Ömer Tarik, et al.
Pubblicazione: (2025)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
di: Xu, Tianyi, et al.
Pubblicazione: (2024)
di: Xu, Tianyi, et al.
Pubblicazione: (2024)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
di: Nguyen, Tuan, et al.
Pubblicazione: (2025)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
di: Ferraz, Thomas Palmeira, et al.
Pubblicazione: (2023)
di: Ferraz, Thomas Palmeira, et al.
Pubblicazione: (2023)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2025)
di: Sudo, Yui, et al.
Pubblicazione: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
di: Peng, Yifan, et al.
Pubblicazione: (2025)
di: Peng, Yifan, et al.
Pubblicazione: (2025)
Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
di: Cuervo, Santiago, et al.
Pubblicazione: (2025)
di: Cuervo, Santiago, et al.
Pubblicazione: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024)
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
di: Goel, Arnav, et al.
Pubblicazione: (2024)
di: Goel, Arnav, et al.
Pubblicazione: (2024)
Audio-to-Score Conversion Model Based on Whisper methodology
di: Zhang, Hongyao, et al.
Pubblicazione: (2024)
di: Zhang, Hongyao, et al.
Pubblicazione: (2024)
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
di: Rao, Rajath, et al.
Pubblicazione: (2025)
di: Rao, Rajath, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Speech foundation models on intelligibility prediction for hearing-impaired listeners
di: Cuervo, Santiago, et al.
Pubblicazione: (2024) -
Scaling Properties of Speech Language Models
di: Cuervo, Santiago, et al.
Pubblicazione: (2024) -
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
di: Wang, Haoyu, et al.
Pubblicazione: (2024) -
Contrastive prediction strategies for unsupervised segmentation and categorization of phonemes and words
di: Cuervo, Santiago, et al.
Pubblicazione: (2021) -
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)