Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration
Fuente:
arXiv
Salvato in:
| Autore principale: | Wang, Haoxuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
di: Nozawa, Kento, et al.
Pubblicazione: (2024)
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
di: Everson, Kevin, et al.
Pubblicazione: (2024)
di: Everson, Kevin, et al.
Pubblicazione: (2024)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
di: Yamashita, Natsuo, et al.
Pubblicazione: (2025)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2025)
Investigating the Impact of Word Informativeness on Speech Emotion Recognition
di: Kakouros, Sofoklis
Pubblicazione: (2025)
di: Kakouros, Sofoklis
Pubblicazione: (2025)
Zero-resource Speech Translation and Recognition with LLMs
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
InterBiasing: Boost Unseen Word Recognition through Biasing Intermediate Predictions
di: Nakagome, Yu, et al.
Pubblicazione: (2024)
di: Nakagome, Yu, et al.
Pubblicazione: (2024)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
di: Liu, Changsong, et al.
Pubblicazione: (2025)
di: Liu, Changsong, et al.
Pubblicazione: (2025)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
di: Jogi, Yash, et al.
Pubblicazione: (2025)
di: Jogi, Yash, et al.
Pubblicazione: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
di: Ao, Junyi, et al.
Pubblicazione: (2024)
di: Ao, Junyi, et al.
Pubblicazione: (2024)
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
di: Wang, Pengcheng, et al.
Pubblicazione: (2025)
di: Wang, Pengcheng, et al.
Pubblicazione: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
di: Zhuo, Le, et al.
Pubblicazione: (2023)
di: Zhuo, Le, et al.
Pubblicazione: (2023)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Whisper Has an Internal Word Aligner
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2025)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2025)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
di: Zhu, Han, et al.
Pubblicazione: (2024)
di: Zhu, Han, et al.
Pubblicazione: (2024)
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
di: Burdisso, Sergio, et al.
Pubblicazione: (2026)
di: Burdisso, Sergio, et al.
Pubblicazione: (2026)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
di: Saliba, Alexandra, et al.
Pubblicazione: (2024)
di: Saliba, Alexandra, et al.
Pubblicazione: (2024)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
di: Wang, Haoyu, et al.
Pubblicazione: (2022)
di: Wang, Haoyu, et al.
Pubblicazione: (2022)
Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper
di: Tran, Hoan My, et al.
Pubblicazione: (2026)
di: Tran, Hoan My, et al.
Pubblicazione: (2026)
Towards few-shot isolated word reading assessment
di: Smit, Reuben, et al.
Pubblicazione: (2025)
di: Smit, Reuben, et al.
Pubblicazione: (2025)
Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge
di: Gao, Miaomiao, et al.
Pubblicazione: (2025)
di: Gao, Miaomiao, et al.
Pubblicazione: (2025)
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study
di: Min, Zeping, et al.
Pubblicazione: (2023)
di: Min, Zeping, et al.
Pubblicazione: (2023)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
di: Visser, Nicol, et al.
Pubblicazione: (2026)
di: Visser, Nicol, et al.
Pubblicazione: (2026)
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
di: Wu, Junkai, et al.
Pubblicazione: (2024)
di: Wu, Junkai, et al.
Pubblicazione: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
di: Higuchi, Yosuke, et al.
Pubblicazione: (2023)
di: Higuchi, Yosuke, et al.
Pubblicazione: (2023)
Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation
di: Ou, Longshen, et al.
Pubblicazione: (2023)
di: Ou, Longshen, et al.
Pubblicazione: (2023)
RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance
di: Chen, Jing-Han, et al.
Pubblicazione: (2026)
di: Chen, Jing-Han, et al.
Pubblicazione: (2026)
Data Augmentation for End-to-end Code-switching Speech Recognition
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
di: Du, Chenpeng, et al.
Pubblicazione: (2020)
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
di: Wang, Guansu, et al.
Pubblicazione: (2025)
di: Wang, Guansu, et al.
Pubblicazione: (2025)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
di: Tee, Hitomi Jin Ling, et al.
Pubblicazione: (2025)
di: Tee, Hitomi Jin Ling, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
di: Nozawa, Kento, et al.
Pubblicazione: (2024) -
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
di: Everson, Kevin, et al.
Pubblicazione: (2024) -
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
di: Yamashita, Natsuo, et al.
Pubblicazione: (2025) -
Investigating the Impact of Word Informativeness on Speech Emotion Recognition
di: Kakouros, Sofoklis
Pubblicazione: (2025) -
Zero-resource Speech Translation and Recognition with LLMs
di: Mundnich, Karel, et al.
Pubblicazione: (2024)