Customizing Speech Recognition Model with Large Language Model Feedback
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ling, Shaoshi, Ye, Guoli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023)
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
von: Nozawa, Kento, et al.
Veröffentlicht: (2024)
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study
von: Min, Zeping, et al.
Veröffentlicht: (2023)
von: Min, Zeping, et al.
Veröffentlicht: (2023)
Benchmarking Automatic Speech Recognition Models for African Languages
von: Nahabwe, Alvin, et al.
Veröffentlicht: (2025)
von: Nahabwe, Alvin, et al.
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
A Deep Learning Automatic Speech Recognition Model for Shona Language
von: Sirora, Leslie Wellington, et al.
Veröffentlicht: (2025)
von: Sirora, Leslie Wellington, et al.
Veröffentlicht: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
Closing the Modality Reasoning Gap for Speech Large Language Models
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Towards Unsupervised Speech Recognition Without Pronunciation Models
von: Ni, Junrui, et al.
Veröffentlicht: (2024)
von: Ni, Junrui, et al.
Veröffentlicht: (2024)
Sequential Editing for Lifelong Training of Speech Recognition Models
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
SEAL: Speech Embedding Alignment Learning for Speech Large Language Model with Retrieval-Augmented Generation
von: Sun, Chunyu, et al.
Veröffentlicht: (2025)
von: Sun, Chunyu, et al.
Veröffentlicht: (2025)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
SpeechVerse: A Large-scale Generalizable Audio Language Model
von: Das, Nilaksh, et al.
Veröffentlicht: (2024)
von: Das, Nilaksh, et al.
Veröffentlicht: (2024)
Investigating Decoder-only Large Language Models for Speech-to-text Translation
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
von: Wang, Siyin, et al.
Veröffentlicht: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
von: Adila, Aulia, et al.
Veröffentlicht: (2024)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2026)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
When Large Language Models Meet Speech: A Survey on Integration Approaches
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hybrid Attention-based Encoder-decoder Model for Efficient Language Model Adaptation
von: Ling, Shaoshi, et al.
Veröffentlicht: (2023) -
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024) -
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024) -
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024) -
Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)