Guardado en:
| Autores principales: | Fang, Zihao, Shen, Yingda, Guan, Zifan, Song, Tongtong, Liu, Zhenyi, Wu, Zhizheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.08046 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
por: Wagner, Dominik, et al.
Publicado: (2023)
por: Wagner, Dominik, et al.
Publicado: (2023)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
por: Zhao, Yiyang, et al.
Publicado: (2024)
por: Zhao, Yiyang, et al.
Publicado: (2024)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
por: Lin, Zhaofeng, et al.
Publicado: (2023)
por: Lin, Zhaofeng, et al.
Publicado: (2023)
Prompting Whisper for Joint Speech Transcription and Diarization
por: Zamyrova, Mariia, et al.
Publicado: (2026)
por: Zamyrova, Mariia, et al.
Publicado: (2026)
Probing Whisper for Dysarthric Speech in Detection and Assessment
por: Yue, Zhengjun, et al.
Publicado: (2025)
por: Yue, Zhengjun, et al.
Publicado: (2025)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2026)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2026)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
por: Boccato, Tommaso, et al.
Publicado: (2026)
por: Boccato, Tommaso, et al.
Publicado: (2026)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
por: Zhou, Haoran, et al.
Publicado: (2025)
por: Zhou, Haoran, et al.
Publicado: (2025)
A Study on Incorporating Whisper for Robust Speech Assessment
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
por: Zhang, Li, et al.
Publicado: (2024)
por: Zhang, Li, et al.
Publicado: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
por: Avdeeva, Anastasia, et al.
Publicado: (2024)
por: Avdeeva, Anastasia, et al.
Publicado: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
por: Guo, Pengcheng, et al.
Publicado: (2024)
por: Guo, Pengcheng, et al.
Publicado: (2024)
Sparsely Shared LoRA on Whisper for Child Speech Recognition
por: Liu, Wei, et al.
Publicado: (2023)
por: Liu, Wei, et al.
Publicado: (2023)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2024)
por: Farhadipour, Aref, et al.
Publicado: (2024)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
por: Zezario, Ryandhimas E., et al.
Publicado: (2025)
por: Zezario, Ryandhimas E., et al.
Publicado: (2025)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2025)
por: Farhadipour, Aref, et al.
Publicado: (2025)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
por: Wang, Shiyao, et al.
Publicado: (2025)
por: Wang, Shiyao, et al.
Publicado: (2025)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
por: Bhuiyan, Mohammed Aman, et al.
Publicado: (2026)
por: Bhuiyan, Mohammed Aman, et al.
Publicado: (2026)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
por: Rouditchenko, Andrew, et al.
Publicado: (2024)
por: Rouditchenko, Andrew, et al.
Publicado: (2024)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
por: Barański, Mateusz, et al.
Publicado: (2025)
por: Barański, Mateusz, et al.
Publicado: (2025)
Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
por: Davudova, Milena, et al.
Publicado: (2025)
por: Davudova, Milena, et al.
Publicado: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
WhisperFlow: speech foundation models in real time
por: Wang, Rongxiang, et al.
Publicado: (2024)
por: Wang, Rongxiang, et al.
Publicado: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
por: Hu, Rui, et al.
Publicado: (2025)
por: Hu, Rui, et al.
Publicado: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
por: Shao, Hang, et al.
Publicado: (2023)
por: Shao, Hang, et al.
Publicado: (2023)
WhisperMask: A Noise Suppressive Mask-Type Microphone for Whisper Speech
por: Hiraki, Hirotaka, et al.
Publicado: (2024)
por: Hiraki, Hirotaka, et al.
Publicado: (2024)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
por: Peng, Yifan, et al.
Publicado: (2025)
por: Peng, Yifan, et al.
Publicado: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
por: Krichli, Tomer, et al.
Publicado: (2025)
por: Krichli, Tomer, et al.
Publicado: (2025)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
por: Sy, Yaya, et al.
Publicado: (2025)
por: Sy, Yaya, et al.
Publicado: (2025)
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
por: Jogi, Yash, et al.
Publicado: (2025)
por: Jogi, Yash, et al.
Publicado: (2025)
Exploiting Music Source Separation for Automatic Lyrics Transcription with Whisper
por: Syed, Jaza, et al.
Publicado: (2025)
por: Syed, Jaza, et al.
Publicado: (2025)
Audio-to-Score Conversion Model Based on Whisper methodology
por: Zhang, Hongyao, et al.
Publicado: (2024)
por: Zhang, Hongyao, et al.
Publicado: (2024)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
por: Jiang, Yicong, et al.
Publicado: (2024)
por: Jiang, Yicong, et al.
Publicado: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
por: Raina, Vyas, et al.
Publicado: (2024)
por: Raina, Vyas, et al.
Publicado: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
por: Kocour, Martin, et al.
Publicado: (2025)
por: Kocour, Martin, et al.
Publicado: (2025)
Ejemplares similares
-
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
por: Wagner, Dominik, et al.
Publicado: (2023) -
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
por: Zhao, Yiyang, et al.
Publicado: (2024) -
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
por: Lin, Zhaofeng, et al.
Publicado: (2023) -
Prompting Whisper for Joint Speech Transcription and Diarization
por: Zamyrova, Mariia, et al.
Publicado: (2026) -
Probing Whisper for Dysarthric Speech in Detection and Assessment
por: Yue, Zhengjun, et al.
Publicado: (2025)