Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
Fuente:
arXiv
Guardado en:
| Autores principales: | Ko, Yuka, Li, Sheng, Yang, Chao-Han Huck, Kawahara, Tatsuya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
por: Sachdev, Rithik, et al.
Publicado: (2024)
por: Sachdev, Rithik, et al.
Publicado: (2024)
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
por: Honda, Tomoki, et al.
Publicado: (2024)
por: Honda, Tomoki, et al.
Publicado: (2024)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2024)
por: Mu, Bingshen, et al.
Publicado: (2024)
Exploration of Adapter for Noise Robust Automatic Speech Recognition
por: Shi, Hao, et al.
Publicado: (2024)
por: Shi, Hao, et al.
Publicado: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)
por: Pusateri, Ernest, et al.
Publicado: (2024)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2025)
por: Mu, Bingshen, et al.
Publicado: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
por: Zheng, Xiuwen, et al.
Publicado: (2026)
por: Zheng, Xiuwen, et al.
Publicado: (2026)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
por: Gao, Yuan, et al.
Publicado: (2025)
por: Gao, Yuan, et al.
Publicado: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
por: Bai, Ye, et al.
Publicado: (2024)
por: Bai, Ye, et al.
Publicado: (2024)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
por: Shi, Hao, et al.
Publicado: (2024)
por: Shi, Hao, et al.
Publicado: (2024)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
Speech Emotion Recognition with ASR Integration
por: Li, Yuanchao
Publicado: (2026)
por: Li, Yuanchao
Publicado: (2026)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
por: Tian, Wenjie, et al.
Publicado: (2026)
por: Tian, Wenjie, et al.
Publicado: (2026)
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
por: Inoue, Koji, et al.
Publicado: (2025)
por: Inoue, Koji, et al.
Publicado: (2025)
Combining Deterministic Enhanced Conditions with Dual-Streaming Encoding for Diffusion-Based Speech Enhancement
por: Shi, Hao, et al.
Publicado: (2025)
por: Shi, Hao, et al.
Publicado: (2025)
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
por: Wei, Linye, et al.
Publicado: (2025)
por: Wei, Linye, et al.
Publicado: (2025)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
por: Yang, Yufeng, et al.
Publicado: (2025)
por: Yang, Yufeng, et al.
Publicado: (2025)
Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses
por: Yang, Yufeng, et al.
Publicado: (2025)
por: Yang, Yufeng, et al.
Publicado: (2025)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
por: Wang, Huimeng, et al.
Publicado: (2024)
por: Wang, Huimeng, et al.
Publicado: (2024)
ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
por: Wei, Victor Junqiu, et al.
Publicado: (2024)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems
por: Robatian, Amin, et al.
Publicado: (2025)
por: Robatian, Amin, et al.
Publicado: (2025)
Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
por: Yang, Chao-Han Huck, et al.
Publicado: (2023)
por: Yang, Chao-Han Huck, et al.
Publicado: (2023)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
por: Aboeitta, Ahmed, et al.
Publicado: (2025)
por: Aboeitta, Ahmed, et al.
Publicado: (2025)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
por: Pulikodan, Sujith, et al.
Publicado: (2025)
por: Pulikodan, Sujith, et al.
Publicado: (2025)
Speech Recognition on TV Series with Video-guided Post-ASR Correction
por: Yang, Haoyuan, et al.
Publicado: (2025)
por: Yang, Haoyuan, et al.
Publicado: (2025)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
por: Zhou, Haoran, et al.
Publicado: (2025)
por: Zhou, Haoran, et al.
Publicado: (2025)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
por: Radhakrishnan, Srijith, et al.
Publicado: (2023)
por: Radhakrishnan, Srijith, et al.
Publicado: (2023)
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems
por: Rai, Anand, et al.
Publicado: (2025)
por: Rai, Anand, et al.
Publicado: (2025)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
por: Wang, He, et al.
Publicado: (2025)
por: Wang, He, et al.
Publicado: (2025)
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
por: Xu, Kai-Tuo, et al.
Publicado: (2025)
por: Xu, Kai-Tuo, et al.
Publicado: (2025)
FUSE: Universal Speech Enhancement using Multi-Stage Fusion of Sparse Compression and Token Generation Models for the URGENT 2025 Challenge
por: Goswami, Nabarun, et al.
Publicado: (2025)
por: Goswami, Nabarun, et al.
Publicado: (2025)
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
por: Xue, Hongfei, et al.
Publicado: (2025)
por: Xue, Hongfei, et al.
Publicado: (2025)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
por: Wang, Weiqing, et al.
Publicado: (2024)
por: Wang, Weiqing, et al.
Publicado: (2024)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
por: Mujtaba, Dena, et al.
Publicado: (2025)
por: Mujtaba, Dena, et al.
Publicado: (2025)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
por: Ochiai, Tsubasa, et al.
Publicado: (2024)
por: Ochiai, Tsubasa, et al.
Publicado: (2024)
EfficientASR: Speech Recognition Network Compression via Attention Redundancy and Chunk-Level FFN Optimization
por: Wang, Jianzong, et al.
Publicado: (2024)
por: Wang, Jianzong, et al.
Publicado: (2024)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
por: Gong, Xun, et al.
Publicado: (2025)
por: Gong, Xun, et al.
Publicado: (2025)
Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
por: Tan, Chao, et al.
Publicado: (2024)
por: Tan, Chao, et al.
Publicado: (2024)
Ejemplares similares
-
Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
por: Sachdev, Rithik, et al.
Publicado: (2024) -
Efficient and Robust Long-Form Speech Recognition with Hybrid H3-Conformer
por: Honda, Tomoki, et al.
Publicado: (2024) -
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2024) -
Exploration of Adapter for Noise Robust Automatic Speech Recognition
por: Shi, Hao, et al.
Publicado: (2024) -
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)