Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Peng, Yang, Yifan, Liang, Zheng, Tan, Tian, Zhang, Shiliang, Chen, Xie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
por: Gong, Xun, et al.
Publicado: (2024)
por: Gong, Xun, et al.
Publicado: (2024)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
por: Zhang, Tian-Hao, et al.
Publicado: (2023)
por: Zhang, Tian-Hao, et al.
Publicado: (2023)
Medical Spoken Named Entity Recognition
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)
por: Pusateri, Ernest, et al.
Publicado: (2024)
Self-Supervised Learning for Multi-Channel Neural Transducer
por: Kojima, Atsushi
Publicado: (2024)
por: Kojima, Atsushi
Publicado: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
por: Shen, Siyuan, et al.
Publicado: (2024)
por: Shen, Siyuan, et al.
Publicado: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
por: Yang, Zijian, et al.
Publicado: (2023)
por: Yang, Zijian, et al.
Publicado: (2023)
Alignment-Free Training for Transducer-based Multi-Talker ASR
por: Moriya, Takafumi, et al.
Publicado: (2024)
por: Moriya, Takafumi, et al.
Publicado: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
por: Cui, Mingyu, et al.
Publicado: (2024)
por: Cui, Mingyu, et al.
Publicado: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
por: Moriya, Takafumi, et al.
Publicado: (2024)
por: Moriya, Takafumi, et al.
Publicado: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
por: Sudo, Yui, et al.
Publicado: (2025)
por: Sudo, Yui, et al.
Publicado: (2025)
Promptformer: Prompted Conformer Transducer for ASR
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
por: Zhao, Zhixian, et al.
Publicado: (2025)
por: Zhao, Zhixian, et al.
Publicado: (2025)
Lightweight Transducer Based on Frame-Level Criterion
por: Wan, Genshun, et al.
Publicado: (2024)
por: Wan, Genshun, et al.
Publicado: (2024)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
por: Du, Chenpeng, et al.
Publicado: (2024)
por: Du, Chenpeng, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
por: Ling, Shaoshi, et al.
Publicado: (2023)
por: Ling, Shaoshi, et al.
Publicado: (2023)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
por: Tang, Yun, et al.
Publicado: (2025)
por: Tang, Yun, et al.
Publicado: (2025)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
por: Xue, Hongfei, et al.
Publicado: (2023)
por: Xue, Hongfei, et al.
Publicado: (2023)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
por: An, Keyu, et al.
Publicado: (2024)
por: An, Keyu, et al.
Publicado: (2024)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
OpusLM: A Family of Open Unified Speech Language Models
por: Tian, Jinchuan, et al.
Publicado: (2025)
por: Tian, Jinchuan, et al.
Publicado: (2025)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
por: Thorbecke, Iuliia, et al.
Publicado: (2024)
por: Thorbecke, Iuliia, et al.
Publicado: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
por: Ma, Ziyang, et al.
Publicado: (2023)
por: Ma, Ziyang, et al.
Publicado: (2023)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
por: Nozawa, Kento, et al.
Publicado: (2024)
por: Nozawa, Kento, et al.
Publicado: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
por: Wang, Yujin, et al.
Publicado: (2022)
por: Wang, Yujin, et al.
Publicado: (2022)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
por: Tian, Jinchuan, et al.
Publicado: (2024)
por: Tian, Jinchuan, et al.
Publicado: (2024)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
por: Tian, Jinchuan, et al.
Publicado: (2025)
por: Tian, Jinchuan, et al.
Publicado: (2025)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
por: Xi, Yu, et al.
Publicado: (2025)
por: Xi, Yu, et al.
Publicado: (2025)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
Customizing Speech Recognition Model with Large Language Model Feedback
por: Ling, Shaoshi, et al.
Publicado: (2025)
por: Ling, Shaoshi, et al.
Publicado: (2025)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
Ejemplares similares
-
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
por: Gong, Xun, et al.
Publicado: (2024) -
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
por: Zhang, Tian-Hao, et al.
Publicado: (2023) -
Medical Spoken Named Entity Recognition
por: Le-Duc, Khai, et al.
Publicado: (2024) -
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
por: Sudo, Yui, et al.
Publicado: (2024) -
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)