Guardado en:
| Autores principales: | Cheng, Gaofeng, Lu, Haitian, Yang, Chengxu, Wang, Xuyang, Li, Ta, Yan, Yonghong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2501.00804 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
por: Fu, Li, et al.
Publicado: (2025)
por: Fu, Li, et al.
Publicado: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
por: Lu, Haitian, et al.
Publicado: (2025)
por: Lu, Haitian, et al.
Publicado: (2025)
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
por: Wang, Ke, et al.
Publicado: (2025)
por: Wang, Ke, et al.
Publicado: (2025)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
por: Liu, Changsong, et al.
Publicado: (2025)
por: Liu, Changsong, et al.
Publicado: (2025)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
por: Huang, Ruizhe, et al.
Publicado: (2024)
por: Huang, Ruizhe, et al.
Publicado: (2024)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
por: Zhu, Han, et al.
Publicado: (2024)
por: Zhu, Han, et al.
Publicado: (2024)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
por: Ren, Bo, et al.
Publicado: (2025)
por: Ren, Bo, et al.
Publicado: (2025)
Improving ASR Contextual Biasing with Guided Attention
por: Tang, Jiyang, et al.
Publicado: (2024)
por: Tang, Jiyang, et al.
Publicado: (2024)
Pronunciation Assessment with Multi-modal Large Language Models
por: Fu, Kaiqi, et al.
Publicado: (2024)
por: Fu, Kaiqi, et al.
Publicado: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2025)
por: Sudo, Yui, et al.
Publicado: (2025)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
por: Gu, Yue, et al.
Publicado: (2025)
por: Gu, Yue, et al.
Publicado: (2025)
Towards Unsupervised Speech Recognition Without Pronunciation Models
por: Ni, Junrui, et al.
Publicado: (2024)
por: Ni, Junrui, et al.
Publicado: (2024)
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
por: Wang, Ke, et al.
Publicado: (2025)
por: Wang, Ke, et al.
Publicado: (2025)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
por: Sun, Guangzhi, et al.
Publicado: (2022)
por: Sun, Guangzhi, et al.
Publicado: (2022)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
por: Nakagome, Yu, et al.
Publicado: (2025)
por: Nakagome, Yu, et al.
Publicado: (2025)
UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
por: Kato, Shuhei
Publicado: (2025)
por: Kato, Shuhei
Publicado: (2025)
Text Injection for Neural Contextual Biasing
por: Meng, Zhong, et al.
Publicado: (2024)
por: Meng, Zhong, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
por: Lin, Zhennan, et al.
Publicado: (2025)
por: Lin, Zhennan, et al.
Publicado: (2025)
Segmentation-free Goodness of Pronunciation
por: Cao, Xinwei, et al.
Publicado: (2025)
por: Cao, Xinwei, et al.
Publicado: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
por: Nguyen, Tuan-Nam, et al.
Publicado: (2025)
por: Nguyen, Tuan-Nam, et al.
Publicado: (2025)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
por: Ellinas, Nikolaos, et al.
Publicado: (2022)
por: Ellinas, Nikolaos, et al.
Publicado: (2022)
A Neural Model for Contextual Biasing Score Learning and Filtering
por: Huang, Wanting, et al.
Publicado: (2025)
por: Huang, Wanting, et al.
Publicado: (2025)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
por: Sun, Siqi, et al.
Publicado: (2024)
por: Sun, Siqi, et al.
Publicado: (2024)
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
InterBiasing: Boost Unseen Word Recognition through Biasing Intermediate Predictions
por: Nakagome, Yu, et al.
Publicado: (2024)
por: Nakagome, Yu, et al.
Publicado: (2024)
Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss
por: Chao, Fu-An, et al.
Publicado: (2025)
por: Chao, Fu-An, et al.
Publicado: (2025)
Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
por: Serditova, Dana, et al.
Publicado: (2025)
por: Serditova, Dana, et al.
Publicado: (2025)
ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
por: Yan, Bi-Cheng, et al.
Publicado: (2024)
por: Yan, Bi-Cheng, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
por: Chen, Yu-Wen, et al.
Publicado: (2023)
por: Chen, Yu-Wen, et al.
Publicado: (2023)
Revisiting Interpolation Augmentation for Speech-to-Text Generation
por: Xu, Chen, et al.
Publicado: (2024)
por: Xu, Chen, et al.
Publicado: (2024)
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech
por: Mujtaba, Dena, et al.
Publicado: (2024)
por: Mujtaba, Dena, et al.
Publicado: (2024)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
por: Li, Shuhe, et al.
Publicado: (2025)
por: Li, Shuhe, et al.
Publicado: (2025)
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
por: Yang, Yifan, et al.
Publicado: (2025)
por: Yang, Yifan, et al.
Publicado: (2025)
Transducer Consistency Regularization for Speech to Text Applications
por: Tseng, Cindy, et al.
Publicado: (2024)
por: Tseng, Cindy, et al.
Publicado: (2024)
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
por: Abdelfattah, Abdullah, et al.
Publicado: (2025)
por: Abdelfattah, Abdullah, et al.
Publicado: (2025)
Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
por: Wang, Siyin, et al.
Publicado: (2024)
por: Wang, Siyin, et al.
Publicado: (2024)
Ejemplares similares
-
PAC: Pronunciation-Aware Contextualized Large Language Model-based Automatic Speech Recognition
por: Fu, Li, et al.
Publicado: (2025) -
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
por: Lu, Haitian, et al.
Publicado: (2025) -
Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment
por: Wang, Ke, et al.
Publicado: (2025) -
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
por: Liu, Changsong, et al.
Publicado: (2025) -
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)