Gespeichert in:
| Hauptverfasser: | Lu, Xugang, Shen, Peng, Tsao, Yu, Kawai, Hisashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.02239 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
von: Lu, Xugang, et al.
Veröffentlicht: (2025)
von: Lu, Xugang, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
von: Shen, Peng, et al.
Veröffentlicht: (2025)
von: Shen, Peng, et al.
Veröffentlicht: (2025)
Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
von: Hira, Medha, et al.
Veröffentlicht: (2024)
von: Hira, Medha, et al.
Veröffentlicht: (2024)
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
von: Yang, Wenhao, et al.
Veröffentlicht: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
von: Wei, Victor Junqiu, et al.
Veröffentlicht: (2024)
AutoMode-ASR: Learning to Select ASR Systems for Better Quality and Cost
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
von: Gündüz, Ahmet, et al.
Veröffentlicht: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
PromptASR for contextualized ASR with controllable style
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
Romanization Encoding For Multilingual ASR
von: Ding, Wen, et al.
Veröffentlicht: (2024)
von: Ding, Wen, et al.
Veröffentlicht: (2024)
HypR: A comprehensive study for ASR hypothesis revising with a reference corpus
von: Wang, Yi-Wei, et al.
Veröffentlicht: (2023)
von: Wang, Yi-Wei, et al.
Veröffentlicht: (2023)
Extending Whisper with prompt tuning to target-speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2023)
von: Ma, Hao, et al.
Veröffentlicht: (2023)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR
von: Chen, Qian, et al.
Veröffentlicht: (2023)
von: Chen, Qian, et al.
Veröffentlicht: (2023)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2024)
Alignment-Free Training for Transducer-based Multi-Talker ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2024)
Towards Rehearsal-Free Multilingual ASR: A LoRA-based Case Study on Whisper
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
von: Xu, Tianyi, et al.
Veröffentlicht: (2024)
Qwen3-ASR Technical Report
von: Shi, Xian, et al.
Veröffentlicht: (2026)
von: Shi, Xian, et al.
Veröffentlicht: (2026)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
von: Tankasala, Srinath, et al.
Veröffentlicht: (2023)
von: Tankasala, Srinath, et al.
Veröffentlicht: (2023)
Diagnostic-Driven Layer-Wise Compensation for Post-Training Quantization of Encoder-Decoder ASR Models
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
Linguistic Knowledge Transfer Learning for Speech Enhancement
von: Hung, Kuo-Hsuan, et al.
Veröffentlicht: (2025)
von: Hung, Kuo-Hsuan, et al.
Veröffentlicht: (2025)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2024)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
von: Ho, Tuan Vu, et al.
Veröffentlicht: (2025)
von: Ho, Tuan Vu, et al.
Veröffentlicht: (2025)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
von: Dutta, Soumya, et al.
Veröffentlicht: (2023)
von: Dutta, Soumya, et al.
Veröffentlicht: (2023)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
von: Tsubaki, Shunsuke, et al.
Veröffentlicht: (2024)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
von: Cooper, Erica, et al.
Veröffentlicht: (2025)
Promptformer: Prompted Conformer Transducer for ASR
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
von: Duarte-Torres, Sergio, et al.
Veröffentlicht: (2024)
Revisiting Acoustic Features for Robust ASR
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
von: Lu, Xugang, et al.
Veröffentlicht: (2025) -
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
von: Shen, Peng, et al.
Veröffentlicht: (2025) -
Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
von: Yang, Wenhao, et al.
Veröffentlicht: (2024) -
CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
von: Hira, Medha, et al.
Veröffentlicht: (2024) -
FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)