Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Chao-Han Huck, Park, Taejin, Gong, Yuan, Li, Yuanchao, Chen, Zhehuai, Lin, Yen-Ting, Chen, Chen, Hu, Yuchen, Dhawan, Kunal, Żelasko, Piotr, Zhang, Chao, Chen, Yun-Nung, Tsao, Yu, Balam, Jagadeesh, Ginsburg, Boris, Siniscalchi, Sabato Marco, Chng, Eng Siong, Bell, Peter, Lai, Catherine, Watanabe, Shinji, Stolcke, Andreas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025)
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025)
Chain-of-Thought Prompting for Speech Translation
von: Hu, Ke, et al.
Veröffentlicht: (2024)
von: Hu, Ke, et al.
Veröffentlicht: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
von: Wang, Weiqing, et al.
Veröffentlicht: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
von: Dhawan, Kunal, et al.
Veröffentlicht: (2024)
von: Dhawan, Kunal, et al.
Veröffentlicht: (2024)
Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
von: Wang, Weiqing, et al.
Veröffentlicht: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2024)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
von: Noroozi, Vahid, et al.
Veröffentlicht: (2024)
von: Noroozi, Vahid, et al.
Veröffentlicht: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
EMMeTT: Efficient Multimodal Machine Translation Training
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Anticipating Future with Large Language Model for Simultaneous Machine Translation
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
von: Noroozi, Vahid, et al.
Veröffentlicht: (2023)
von: Noroozi, Vahid, et al.
Veröffentlicht: (2023)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2024)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend
von: Jukić, Ante, et al.
Veröffentlicht: (2024)
von: Jukić, Ante, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
von: Wan, Zhen, et al.
Veröffentlicht: (2026)
von: Wan, Zhen, et al.
Veröffentlicht: (2026)
Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission
von: Thakur, Nirmalya Mallick, et al.
Veröffentlicht: (2025)
von: Thakur, Nirmalya Mallick, et al.
Veröffentlicht: (2025)
Joint Tensor-Train Parameterization for Efficient and Expressive Low-Rank Adaptation
von: Qi, Jun, et al.
Veröffentlicht: (2025)
von: Qi, Jun, et al.
Veröffentlicht: (2025)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
von: Peng, Yifan, et al.
Veröffentlicht: (2024) -
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025) -
Chain-of-Thought Prompting for Speech Translation
von: Hu, Ke, et al.
Veröffentlicht: (2024) -
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024) -
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)