Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Jiahui, Shi, Hao, Cui, Chenrui, Wang, Tianrui, Liu, Hexin, Ni, Zhaoheng, Ye, Lingxuan, Wang, Longbiao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
por: Cui, Zhongjian, et al.
Publicado: (2025)
por: Cui, Zhongjian, et al.
Publicado: (2025)
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
por: Wang, Junyu, et al.
Publicado: (2025)
por: Wang, Junyu, et al.
Publicado: (2025)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
por: Shi, Hao, et al.
Publicado: (2024)
por: Shi, Hao, et al.
Publicado: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
por: Zhou, Haoran, et al.
Publicado: (2025)
por: Zhou, Haoran, et al.
Publicado: (2025)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
por: Zhang, Li, et al.
Publicado: (2024)
por: Zhang, Li, et al.
Publicado: (2024)
Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
por: Yang, Hongli, et al.
Publicado: (2025)
por: Yang, Hongli, et al.
Publicado: (2025)
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement
por: Zhang, Qiquan, et al.
Publicado: (2024)
por: Zhang, Qiquan, et al.
Publicado: (2024)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
por: Wang, Junyu, et al.
Publicado: (2024)
por: Wang, Junyu, et al.
Publicado: (2024)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
por: Zezario, Ryandhimas E., et al.
Publicado: (2025)
por: Zezario, Ryandhimas E., et al.
Publicado: (2025)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
por: Boccato, Tommaso, et al.
Publicado: (2026)
por: Boccato, Tommaso, et al.
Publicado: (2026)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
por: Wang, Tianrui, et al.
Publicado: (2024)
por: Wang, Tianrui, et al.
Publicado: (2024)
Deepfake Detection of Singing Voices With Whisper Encodings
por: Sharma, Falguni, et al.
Publicado: (2025)
por: Sharma, Falguni, et al.
Publicado: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
por: Zhao, Yiyang, et al.
Publicado: (2024)
por: Zhao, Yiyang, et al.
Publicado: (2024)
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
por: Wang, Junyu, et al.
Publicado: (2025)
por: Wang, Junyu, et al.
Publicado: (2025)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
Whispy: Adapting STT Whisper Models to Real-Time Environments
por: Bevilacqua, Antonio, et al.
Publicado: (2024)
por: Bevilacqua, Antonio, et al.
Publicado: (2024)
WhisperFlow: speech foundation models in real time
por: Wang, Rongxiang, et al.
Publicado: (2024)
por: Wang, Rongxiang, et al.
Publicado: (2024)
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
por: Mu, Bingshen, et al.
Publicado: (2026)
por: Mu, Bingshen, et al.
Publicado: (2026)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
por: Guo, Pengcheng, et al.
Publicado: (2024)
por: Guo, Pengcheng, et al.
Publicado: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
por: Wang, Junyu, et al.
Publicado: (2025)
por: Wang, Junyu, et al.
Publicado: (2025)
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
por: Zezario, Ryandhimas E., et al.
Publicado: (2023)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
por: Kocour, Martin, et al.
Publicado: (2025)
por: Kocour, Martin, et al.
Publicado: (2025)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
por: Shu, Yuchun, et al.
Publicado: (2024)
por: Shu, Yuchun, et al.
Publicado: (2024)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
por: Wang, Shiyao, et al.
Publicado: (2025)
por: Wang, Shiyao, et al.
Publicado: (2025)
Oral Tradition-Encoded NanyinHGNN: Integrating Nanyin Music Preservation and Generation through a Pipa-Centric Dataset
por: Xiahou, Jianbing, et al.
Publicado: (2025)
por: Xiahou, Jianbing, et al.
Publicado: (2025)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
por: Fang, Zihao, et al.
Publicado: (2026)
por: Fang, Zihao, et al.
Publicado: (2026)
Adapting Language Balance in Code-Switching Speech
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
Prompting Whisper for Joint Speech Transcription and Diarization
por: Zamyrova, Mariia, et al.
Publicado: (2026)
por: Zamyrova, Mariia, et al.
Publicado: (2026)
Probing Whisper for Dysarthric Speech in Detection and Assessment
por: Yue, Zhengjun, et al.
Publicado: (2025)
por: Yue, Zhengjun, et al.
Publicado: (2025)
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
por: Hu, Rui, et al.
Publicado: (2025)
por: Hu, Rui, et al.
Publicado: (2025)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
por: Zhao, Junchuan, et al.
Publicado: (2026)
por: Zhao, Junchuan, et al.
Publicado: (2026)
Attention-Guided Adaptation for Code-Switching Speech Recognition
por: Aditya, Bobbi, et al.
Publicado: (2023)
por: Aditya, Bobbi, et al.
Publicado: (2023)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
por: Ma, Guobin, et al.
Publicado: (2026)
por: Ma, Guobin, et al.
Publicado: (2026)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
por: Gong, Cheng, et al.
Publicado: (2023)
por: Gong, Cheng, et al.
Publicado: (2023)
Ejemplares similares
-
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
por: Cui, Zhongjian, et al.
Publicado: (2025) -
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
por: Wang, Junyu, et al.
Publicado: (2025) -
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
por: Shi, Hao, et al.
Publicado: (2024) -
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
por: Zhou, Haoran, et al.
Publicado: (2025) -
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
por: Zhang, Li, et al.
Publicado: (2024)