Retrieval-Augmented Self-Taught Reasoning Model with Adaptive Chain-of-Thought for ASR Named Entity Correction
Fuente:
arXiv
Guardado en:
| Autores principales: | An, Junjie, Tian, Jingguang, Wang, Tianyi, Gao, Yu, Mou, Xiaofeng, Xu, Yi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
por: Feng, Fuyuan, et al.
Publicado: (2026)
por: Feng, Fuyuan, et al.
Publicado: (2026)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)
por: Pusateri, Ernest, et al.
Publicado: (2024)
End-to-End Direction-Aware Keyword Spotting with Spatial Priors in Noisy Environments
por: Wang, Rui, et al.
Publicado: (2026)
por: Wang, Rui, et al.
Publicado: (2026)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
por: Xi, Yu, et al.
Publicado: (2024)
por: Xi, Yu, et al.
Publicado: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
por: Sun, Haoqin, et al.
Publicado: (2025)
por: Sun, Haoqin, et al.
Publicado: (2025)
Retrieval Augmented Generation based context discovery for ASR
por: Siskos, Dimitrios, et al.
Publicado: (2025)
por: Siskos, Dimitrios, et al.
Publicado: (2025)
Discrete Audio Representations for Automated Audio Captioning
por: Tian, Jingguang, et al.
Publicado: (2025)
por: Tian, Jingguang, et al.
Publicado: (2025)
Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning
por: Chang, Chih-Cheng, et al.
Publicado: (2025)
por: Chang, Chih-Cheng, et al.
Publicado: (2025)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning
por: Kong, YuXiang, et al.
Publicado: (2025)
por: Kong, YuXiang, et al.
Publicado: (2025)
MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
Interpretable Audio Editing Evaluation via Chain-of-Thought Difference-Commonality Reasoning with Multimodal LLMs
por: Jia, Yuhang, et al.
Publicado: (2025)
por: Jia, Yuhang, et al.
Publicado: (2025)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
Romanization Encoding For Multilingual ASR
por: Ding, Wen, et al.
Publicado: (2024)
por: Ding, Wen, et al.
Publicado: (2024)
persoDA: Personalized Data Augmentation for Personalized ASR
por: Parada, Pablo Peso, et al.
Publicado: (2025)
por: Parada, Pablo Peso, et al.
Publicado: (2025)
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
por: Saon, George, et al.
Publicado: (2026)
por: Saon, George, et al.
Publicado: (2026)
Multi-Channel Multi-Speaker ASR Using Target Speaker's Solo Segment
por: Shao, Yiwen, et al.
Publicado: (2024)
por: Shao, Yiwen, et al.
Publicado: (2024)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
por: Zhao, Qiuming, et al.
Publicado: (2024)
por: Zhao, Qiuming, et al.
Publicado: (2024)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
por: Gong, Xun, et al.
Publicado: (2025)
por: Gong, Xun, et al.
Publicado: (2025)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
por: Li, Shaojun, et al.
Publicado: (2024)
por: Li, Shaojun, et al.
Publicado: (2024)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
por: Lee, Chia-Yu, et al.
Publicado: (2026)
por: Lee, Chia-Yu, et al.
Publicado: (2026)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
por: Li, Yuang, et al.
Publicado: (2024)
por: Li, Yuang, et al.
Publicado: (2024)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
por: Li, Li, et al.
Publicado: (2026)
por: Li, Li, et al.
Publicado: (2026)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
por: Wang, Peng, et al.
Publicado: (2023)
por: Wang, Peng, et al.
Publicado: (2023)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
por: Tian, Jingguang, et al.
Publicado: (2024)
por: Tian, Jingguang, et al.
Publicado: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
por: Zhang, Jisi, et al.
Publicado: (2024)
por: Zhang, Jisi, et al.
Publicado: (2024)
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
por: Singh, Vishwanath Pratap, et al.
Publicado: (2024)
por: Singh, Vishwanath Pratap, et al.
Publicado: (2024)
MDM-ASR: Bridging Accuracy and Efficiency in ASR with Diffusion-Based Non-Autoregressive Decoding
por: Yen, Hao, et al.
Publicado: (2026)
por: Yen, Hao, et al.
Publicado: (2026)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
por: Zhou, Jiaming, et al.
Publicado: (2023)
por: Zhou, Jiaming, et al.
Publicado: (2023)
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
por: Xiang, Yang, et al.
Publicado: (2025)
por: Xiang, Yang, et al.
Publicado: (2025)
Large Language Models based ASR Error Correction for Child Conversations
por: Xu, Anfeng, et al.
Publicado: (2025)
por: Xu, Anfeng, et al.
Publicado: (2025)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
por: Kim, Jongsuk, et al.
Publicado: (2025)
por: Kim, Jongsuk, et al.
Publicado: (2025)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
por: Moriya, Takafumi, et al.
Publicado: (2025)
por: Moriya, Takafumi, et al.
Publicado: (2025)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
por: Wang, Huimeng, et al.
Publicado: (2024)
por: Wang, Huimeng, et al.
Publicado: (2024)
Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions
por: Altinok, Duygu
Publicado: (2025)
por: Altinok, Duygu
Publicado: (2025)
Ejemplares similares
-
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
por: Feng, Fuyuan, et al.
Publicado: (2026) -
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024) -
End-to-End Direction-Aware Keyword Spotting with Spatial Priors in Noisy Environments
por: Wang, Rui, et al.
Publicado: (2026) -
Learning Emotion-Invariant Speaker Representations for Speaker Verification
por: Tian, Jingguang, et al.
Publicado: (2025) -
Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
por: Xi, Yu, et al.
Publicado: (2024)