UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
Fuente:
arXiv
Salvato in:
| Autori principali: | Guo, Jiaxin, Wang, Minghan, Qiao, Xiaosong, Wei, Daimeng, Shang, Hengchao, Li, Zongyao, Yu, Zhengzhe, Li, Yinglu, Su, Chang, Zhang, Min, Tao, Shimin, Yang, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
An End-to-End Speech Summarization Using Large Language Model
di: Shang, Hengchao, et al.
Pubblicazione: (2024)
di: Shang, Hengchao, et al.
Pubblicazione: (2024)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
di: Pusateri, Ernest, et al.
Pubblicazione: (2024)
di: Pusateri, Ernest, et al.
Pubblicazione: (2024)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
di: Shu, Yuchun, et al.
Pubblicazione: (2024)
di: Shu, Yuchun, et al.
Pubblicazione: (2024)
Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
di: Tang, Zhiyuan, et al.
Pubblicazione: (2024)
di: Tang, Zhiyuan, et al.
Pubblicazione: (2024)
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
di: Li, Yuang, et al.
Pubblicazione: (2024)
di: Li, Yuang, et al.
Pubblicazione: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
di: Yang, Da-Hee, et al.
Pubblicazione: (2026)
di: Yang, Da-Hee, et al.
Pubblicazione: (2026)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
di: Mu, Bingshen, et al.
Pubblicazione: (2024)
di: Mu, Bingshen, et al.
Pubblicazione: (2024)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
di: Lin, Yuke, et al.
Pubblicazione: (2025)
di: Lin, Yuke, et al.
Pubblicazione: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning
di: Chang, Chih-Cheng, et al.
Pubblicazione: (2025)
di: Chang, Chih-Cheng, et al.
Pubblicazione: (2025)
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems
di: Robatian, Amin, et al.
Pubblicazione: (2025)
di: Robatian, Amin, et al.
Pubblicazione: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models
di: Frieske, Rita, et al.
Pubblicazione: (2024)
di: Frieske, Rita, et al.
Pubblicazione: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
di: Wei, Linye, et al.
Pubblicazione: (2025)
di: Wei, Linye, et al.
Pubblicazione: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
di: Xue, Hongfei, et al.
Pubblicazione: (2024)
Augmenting Polish Automatic Speech Recognition System With Synthetic Data
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2024)
di: Bondaruk, Łukasz, et al.
Pubblicazione: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
di: Farhadipour, Aref, et al.
Pubblicazione: (2024)
di: Farhadipour, Aref, et al.
Pubblicazione: (2024)
Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2023)
di: Eeckt, Steven Vander, et al.
Pubblicazione: (2023)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
Dynamic Data Pruning for Automatic Speech Recognition
di: Xiao, Qiao, et al.
Pubblicazione: (2024)
di: Xiao, Qiao, et al.
Pubblicazione: (2024)
Unsupervised Single-Channel Speech Separation with a Diffusion Prior under Speaker-Embedding Guidance
di: Shi, Runwu, et al.
Pubblicazione: (2025)
di: Shi, Runwu, et al.
Pubblicazione: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
di: Chen, Peikun, et al.
Pubblicazione: (2024)
di: Chen, Peikun, et al.
Pubblicazione: (2024)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
di: Poncelet, Jakob, et al.
Pubblicazione: (2025)
di: Poncelet, Jakob, et al.
Pubblicazione: (2025)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
di: Zheng, Xiuwen, et al.
Pubblicazione: (2026)
di: Zheng, Xiuwen, et al.
Pubblicazione: (2026)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
di: Dai, Yuhang, et al.
Pubblicazione: (2025)
di: Dai, Yuhang, et al.
Pubblicazione: (2025)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Towards Unsupervised Speech Recognition Without Pronunciation Models
di: Ni, Junrui, et al.
Pubblicazione: (2024)
di: Ni, Junrui, et al.
Pubblicazione: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
Reducing Geographic Disparities in Automatic Speech Recognition via Elastic Weight Consolidation
di: Trinh, Viet Anh, et al.
Pubblicazione: (2022)
di: Trinh, Viet Anh, et al.
Pubblicazione: (2022)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
di: Ko, Yuka, et al.
Pubblicazione: (2024)
di: Ko, Yuka, et al.
Pubblicazione: (2024)
Documenti analoghi
-
An End-to-End Speech Summarization Using Large Language Model
di: Shang, Hengchao, et al.
Pubblicazione: (2024) -
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
di: Li, Shaojun, et al.
Pubblicazione: (2024) -
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
di: Li, Shaojun, et al.
Pubblicazione: (2024) -
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
di: Pusateri, Ernest, et al.
Pubblicazione: (2024) -
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
di: Shu, Yuchun, et al.
Pubblicazione: (2024)