Extending Whisper with prompt tuning to target-speaker ASR
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ma, Hao, Peng, Zhiyuan, Shao, Mingjie, Li, Jing, Liu, Ju |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
par: Thorbecke, Iuliia, et autres
Publié: (2024)
par: Thorbecke, Iuliia, et autres
Publié: (2024)
Quantizing Whisper-small: How design choices affect ASR performance
par: Söhler, Arthur, et autres
Publié: (2025)
par: Söhler, Arthur, et autres
Publié: (2025)
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
par: Orhon, Atila, et autres
Publié: (2025)
par: Orhon, Atila, et autres
Publié: (2025)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
par: Özyilmaz, Ömer Tarik, et autres
Publié: (2025)
par: Özyilmaz, Ömer Tarik, et autres
Publié: (2025)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
par: Xu, Tianyi, et autres
Publié: (2024)
par: Xu, Tianyi, et autres
Publié: (2024)
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish
par: Cumbal, Ronald, et autres
Publié: (2024)
par: Cumbal, Ronald, et autres
Publié: (2024)
Hierarchical speaker representation for target speaker extraction
par: He, Shulin, et autres
Publié: (2022)
par: He, Shulin, et autres
Publié: (2022)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
par: Shekoufandeh, Golshid, et autres
Publié: (2025)
par: Shekoufandeh, Golshid, et autres
Publié: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
par: Anidjar, Or Haim, et autres
Publié: (2024)
par: Anidjar, Or Haim, et autres
Publié: (2024)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
par: Wang, Haoyu, et autres
Publié: (2024)
par: Wang, Haoyu, et autres
Publié: (2024)
Language-Queried Target Sound Extraction Without Parallel Training Data
par: Ma, Hao, et autres
Publié: (2024)
par: Ma, Hao, et autres
Publié: (2024)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
par: Shao, Hang, et autres
Publié: (2023)
par: Shao, Hang, et autres
Publié: (2023)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
par: Grossman, Raymond, et autres
Publié: (2025)
par: Grossman, Raymond, et autres
Publié: (2025)
Target Speaker ASR with Whisper
par: Polok, Alexander, et autres
Publié: (2024)
par: Polok, Alexander, et autres
Publié: (2024)
Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages
par: Liang, Siyu, et autres
Publié: (2025)
par: Liang, Siyu, et autres
Publié: (2025)
Improving curriculum learning for target speaker extraction with synthetic speakers
par: Liu, Yun, et autres
Publié: (2024)
par: Liu, Yun, et autres
Publié: (2024)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
par: Song, Zheshu, et autres
Publié: (2024)
par: Song, Zheshu, et autres
Publié: (2024)
Efficient Multilingual ASR Finetuning via LoRA Language Experts
par: Li, Jiahong, et autres
Publié: (2025)
par: Li, Jiahong, et autres
Publié: (2025)
A Benchmark for Multi-speaker Anonymization
par: Miao, Xiaoxiao, et autres
Publié: (2024)
par: Miao, Xiaoxiao, et autres
Publié: (2024)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
par: Attia, Ahmed Adel, et autres
Publié: (2023)
par: Attia, Ahmed Adel, et autres
Publié: (2023)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
par: Zhao, Jiahui, et autres
Publié: (2024)
par: Zhao, Jiahui, et autres
Publié: (2024)
PromptASR for contextualized ASR with controllable style
par: Yang, Xiaoyu, et autres
Publié: (2023)
par: Yang, Xiaoyu, et autres
Publié: (2023)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
par: Ma, Yingyi, et autres
Publié: (2024)
par: Ma, Yingyi, et autres
Publié: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
par: Raina, Vyas, et autres
Publié: (2024)
par: Raina, Vyas, et autres
Publié: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
par: Guo, Pengcheng, et autres
Publié: (2024)
par: Guo, Pengcheng, et autres
Publié: (2024)
The THUEE System Description for the IARPA OpenASR21 Challenge
par: Zhao, Jing, et autres
Publié: (2022)
par: Zhao, Jing, et autres
Publié: (2022)
An approach to optimize inference of the DIART speaker diarization pipeline
par: Aperdannier, Roman, et autres
Publié: (2024)
par: Aperdannier, Roman, et autres
Publié: (2024)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
par: Sivasankaran, Sunit, et autres
Publié: (2024)
par: Sivasankaran, Sunit, et autres
Publié: (2024)
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
par: Geng, Xuelong, et autres
Publié: (2024)
par: Geng, Xuelong, et autres
Publié: (2024)
ASR Error Correction using Large Language Models
par: Ma, Rao, et autres
Publié: (2024)
par: Ma, Rao, et autres
Publié: (2024)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
par: Shakeel, Muhammad, et autres
Publié: (2025)
par: Shakeel, Muhammad, et autres
Publié: (2025)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
par: Hu, Rui, et autres
Publié: (2025)
par: Hu, Rui, et autres
Publié: (2025)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
par: Zhuo, Le, et autres
Publié: (2023)
par: Zhuo, Le, et autres
Publié: (2023)
Qwen3-ASR Technical Report
par: Shi, Xian, et autres
Publié: (2026)
par: Shi, Xian, et autres
Publié: (2026)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
par: Li, Shaojun, et autres
Publié: (2024)
par: Li, Shaojun, et autres
Publié: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
par: Nguyen, Thai-Binh, et autres
Publié: (2024)
par: Nguyen, Thai-Binh, et autres
Publié: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
par: Xie, Yuan, et autres
Publié: (2026)
par: Xie, Yuan, et autres
Publié: (2026)
Mamba for Streaming ASR Combined with Unimodal Aggregation
par: Fang, Ying, et autres
Publié: (2024)
par: Fang, Ying, et autres
Publié: (2024)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
par: Li, Yuang, et autres
Publié: (2024)
par: Li, Yuang, et autres
Publié: (2024)
Can Whisper perform speech-based in-context learning?
par: Wang, Siyin, et autres
Publié: (2023)
par: Wang, Siyin, et autres
Publié: (2023)
Documents similaires
-
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
par: Thorbecke, Iuliia, et autres
Publié: (2024) -
Quantizing Whisper-small: How design choices affect ASR performance
par: Söhler, Arthur, et autres
Publié: (2025) -
WhisperKit: On-device Real-time ASR with Billion-Scale Transformers
par: Orhon, Atila, et autres
Publié: (2025) -
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
par: Özyilmaz, Ömer Tarik, et autres
Publié: (2025) -
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
par: Xu, Tianyi, et autres
Publié: (2024)