Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration
Fuente:
arXiv
Guardado en:
| Autor principal: | Wang, Haoxuan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
por: Nozawa, Kento, et al.
Publicado: (2024)
por: Nozawa, Kento, et al.
Publicado: (2024)
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
por: Everson, Kevin, et al.
Publicado: (2024)
por: Everson, Kevin, et al.
Publicado: (2024)
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
por: Yamashita, Natsuo, et al.
Publicado: (2025)
por: Yamashita, Natsuo, et al.
Publicado: (2025)
Investigating the Impact of Word Informativeness on Speech Emotion Recognition
por: Kakouros, Sofoklis
Publicado: (2025)
por: Kakouros, Sofoklis
Publicado: (2025)
Zero-resource Speech Translation and Recognition with LLMs
por: Mundnich, Karel, et al.
Publicado: (2024)
por: Mundnich, Karel, et al.
Publicado: (2024)
InterBiasing: Boost Unseen Word Recognition through Biasing Intermediate Predictions
por: Nakagome, Yu, et al.
Publicado: (2024)
por: Nakagome, Yu, et al.
Publicado: (2024)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
por: Liu, Changsong, et al.
Publicado: (2025)
por: Liu, Changsong, et al.
Publicado: (2025)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
por: Chen, Chen, et al.
Publicado: (2024)
por: Chen, Chen, et al.
Publicado: (2024)
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Improving Rare-Word Recognition of Whisper in Zero-Shot Settings
por: Jogi, Yash, et al.
Publicado: (2025)
por: Jogi, Yash, et al.
Publicado: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
por: Hu, Ke, et al.
Publicado: (2025)
por: Hu, Ke, et al.
Publicado: (2025)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
por: Ao, Junyi, et al.
Publicado: (2024)
por: Ao, Junyi, et al.
Publicado: (2024)
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
por: Wang, Pengcheng, et al.
Publicado: (2025)
por: Wang, Pengcheng, et al.
Publicado: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
por: Park, Chanho, et al.
Publicado: (2024)
por: Park, Chanho, et al.
Publicado: (2024)
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT
por: Zhuo, Le, et al.
Publicado: (2023)
por: Zhuo, Le, et al.
Publicado: (2023)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2024)
por: Deng, Keqi, et al.
Publicado: (2024)
Whisper Has an Internal Word Aligner
por: Yeh, Sung-Lin, et al.
Publicado: (2025)
por: Yeh, Sung-Lin, et al.
Publicado: (2025)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
por: Zhu, Han, et al.
Publicado: (2024)
por: Zhu, Han, et al.
Publicado: (2024)
Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
por: Burdisso, Sergio, et al.
Publicado: (2026)
por: Burdisso, Sergio, et al.
Publicado: (2026)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
por: Saliba, Alexandra, et al.
Publicado: (2024)
por: Saliba, Alexandra, et al.
Publicado: (2024)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
por: Wang, Haoyu, et al.
Publicado: (2022)
por: Wang, Haoyu, et al.
Publicado: (2022)
Deepfake Word Detection by Next-token Prediction using Fine-tuned Whisper
por: Tran, Hoan My, et al.
Publicado: (2026)
por: Tran, Hoan My, et al.
Publicado: (2026)
Towards few-shot isolated word reading assessment
por: Smit, Reuben, et al.
Publicado: (2025)
por: Smit, Reuben, et al.
Publicado: (2025)
Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge
por: Gao, Miaomiao, et al.
Publicado: (2025)
por: Gao, Miaomiao, et al.
Publicado: (2025)
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study
por: Min, Zeping, et al.
Publicado: (2023)
por: Min, Zeping, et al.
Publicado: (2023)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
por: Yen, Hao, et al.
Publicado: (2024)
por: Yen, Hao, et al.
Publicado: (2024)
Visually grounded few-shot word learning in low-resource settings
por: Nortje, Leanne, et al.
Publicado: (2023)
por: Nortje, Leanne, et al.
Publicado: (2023)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
por: Hu, Jiliang, et al.
Publicado: (2025)
por: Hu, Jiliang, et al.
Publicado: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
por: Visser, Nicol, et al.
Publicado: (2026)
por: Visser, Nicol, et al.
Publicado: (2026)
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
por: Wu, Junkai, et al.
Publicado: (2024)
por: Wu, Junkai, et al.
Publicado: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
por: Fujita, Kenichi, et al.
Publicado: (2024)
por: Fujita, Kenichi, et al.
Publicado: (2024)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
por: von Neumann, Thilo, et al.
Publicado: (2023)
por: von Neumann, Thilo, et al.
Publicado: (2023)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
por: Higuchi, Yosuke, et al.
Publicado: (2023)
por: Higuchi, Yosuke, et al.
Publicado: (2023)
Joint Learning of Wording and Formatting for Singable Melody-to-Lyric Generation
por: Ou, Longshen, et al.
Publicado: (2023)
por: Ou, Longshen, et al.
Publicado: (2023)
RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance
por: Chen, Jing-Han, et al.
Publicado: (2026)
por: Chen, Jing-Han, et al.
Publicado: (2026)
Data Augmentation for End-to-end Code-switching Speech Recognition
por: Du, Chenpeng, et al.
Publicado: (2020)
por: Du, Chenpeng, et al.
Publicado: (2020)
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
por: Parcollet, Titouan, et al.
Publicado: (2023)
por: Parcollet, Titouan, et al.
Publicado: (2023)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
por: Wang, Guansu, et al.
Publicado: (2025)
por: Wang, Guansu, et al.
Publicado: (2025)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
por: Tee, Hitomi Jin Ling, et al.
Publicado: (2025)
por: Tee, Hitomi Jin Ling, et al.
Publicado: (2025)
Ejemplares similares
-
Enhancing Large Language Model-based Speech Recognition by Contextualization for Rare and Ambiguous Words
por: Nozawa, Kento, et al.
Publicado: (2024) -
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
por: Everson, Kevin, et al.
Publicado: (2024) -
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
por: Yamashita, Natsuo, et al.
Publicado: (2025) -
Investigating the Impact of Word Informativeness on Speech Emotion Recognition
por: Kakouros, Sofoklis
Publicado: (2025) -
Zero-resource Speech Translation and Recognition with LLMs
por: Mundnich, Karel, et al.
Publicado: (2024)