Multi-Modal Retrieval For Large Language Model Based Speech Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Kolehmainen, Jari, Gourav, Aditya, Shivakumar, Prashanth Gurunath, Gu, Yile, Gandhe, Ankur, Rastrow, Ariya, Strimel, Grant, Bulyko, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Speech Recognition Rescoring with Large Speech-Text Foundation Models
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024)
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024)
Group Relative Policy Optimization for Speech Recognition
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2025)
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2025)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
by: Pandey, Rahul, et al.
Published: (2023)
by: Pandey, Rahul, et al.
Published: (2023)
Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
by: Lin, Guan-Ting, et al.
Published: (2024)
by: Lin, Guan-Ting, et al.
Published: (2024)
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
by: Everson, Kevin, et al.
Published: (2024)
by: Everson, Kevin, et al.
Published: (2024)
Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
by: Yu, Yu, et al.
Published: (2023)
by: Yu, Yu, et al.
Published: (2023)
Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition
by: Yu, Yu, et al.
Published: (2024)
by: Yu, Yu, et al.
Published: (2024)
Streaming Speech-to-Confusion Network Speech Recognition
by: Filimonov, Denis, et al.
Published: (2023)
by: Filimonov, Denis, et al.
Published: (2023)
Phone Duration Modeling for Speaker Age Estimation in Children
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2021)
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2021)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
by: Lin, Guan-Ting, et al.
Published: (2023)
by: Lin, Guan-Ting, et al.
Published: (2023)
Two-pass Endpoint Detection for Speech Recognition
by: Raju, Anirudh, et al.
Published: (2024)
by: Raju, Anirudh, et al.
Published: (2024)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
by: Yang, Chao-Han Huck, et al.
Published: (2023)
by: Yang, Chao-Han Huck, et al.
Published: (2023)
Application of Audio Fingerprinting Techniques for Real-Time Scalable Speech Retrieval and Speech Clusterization
by: Altwlkany, Kemal, et al.
Published: (2024)
by: Altwlkany, Kemal, et al.
Published: (2024)
FusID: Modality-Fused Semantic IDs for Generative Music Recommendation
by: Kim, Haven, et al.
Published: (2026)
by: Kim, Haven, et al.
Published: (2026)
Language-based Audio Retrieval with Co-Attention Networks
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
by: Abootorabi, Mohammad Mahdi, et al.
Published: (2024)
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
by: Zhang, Dengming, et al.
Published: (2024)
by: Zhang, Dengming, et al.
Published: (2024)
Expressivity-aware Music Performance Retrieval using Mid-level Perceptual Features and Emotion Word Embeddings
by: Chowdhury, Shreyan, et al.
Published: (2024)
by: Chowdhury, Shreyan, et al.
Published: (2024)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
by: Lin, Chyi-Jiunn, et al.
Published: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
by: Lu, Yen-Ju, et al.
Published: (2024)
by: Lu, Yen-Ju, et al.
Published: (2024)
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
Diff4Steer: Steerable Diffusion Prior for Generative Music Retrieval with Semantic Guidance
by: Bao, Xuchan, et al.
Published: (2024)
by: Bao, Xuchan, et al.
Published: (2024)
JEPOO: Highly Accurate Joint Estimation of Pitch, Onset and Offset for Music Information Retrieval
by: Wei, Haojie, et al.
Published: (2023)
by: Wei, Haojie, et al.
Published: (2023)
Attention-Guided Adaptation for Code-Switching Speech Recognition
by: Aditya, Bobbi, et al.
Published: (2023)
by: Aditya, Bobbi, et al.
Published: (2023)
SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech
by: Sabra, Adam, et al.
Published: (2024)
by: Sabra, Adam, et al.
Published: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
by: Pusateri, Ernest, et al.
Published: (2024)
by: Pusateri, Ernest, et al.
Published: (2024)
Pretrained Conformers for Audio Fingerprinting and Retrieval
by: Altwlkany, Kemal, et al.
Published: (2025)
by: Altwlkany, Kemal, et al.
Published: (2025)
Dissecting Temporal Understanding in Text-to-Audio Retrieval
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
by: Oncescu, Andreea-Maria, et al.
Published: (2024)
Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
by: Loweimi, Erfan, et al.
Published: (2025)
by: Loweimi, Erfan, et al.
Published: (2025)
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering
by: Rackauckas, Zackary, et al.
Published: (2025)
by: Rackauckas, Zackary, et al.
Published: (2025)
Exploring Diverse Sounds: Identifying Outliers in a Music Corpus
by: Cai, Le, et al.
Published: (2024)
by: Cai, Le, et al.
Published: (2024)
Music Discovery Dialogue Generation Using Human Intent Analysis and Large Language Models
by: Doh, SeungHeon, et al.
Published: (2024)
by: Doh, SeungHeon, et al.
Published: (2024)
Track Role Prediction of Single-Instrumental Sequences
by: Han, Changheon, et al.
Published: (2024)
by: Han, Changheon, et al.
Published: (2024)
LARP: Language Audio Relational Pre-training for Cold-Start Playlist Continuation
by: Salganik, Rebecca, et al.
Published: (2024)
by: Salganik, Rebecca, et al.
Published: (2024)
Exploring GPT's Ability as a Judge in Music Understanding
by: Fang, Kun, et al.
Published: (2025)
by: Fang, Kun, et al.
Published: (2025)
Multi-Sample Dynamic Time Warping for Few-Shot Keyword Spotting
by: Wilkinghoff, Kevin, et al.
Published: (2024)
by: Wilkinghoff, Kevin, et al.
Published: (2024)
Do Captioning Metrics Reflect Music Semantic Alignment?
by: Lee, Jinwoo, et al.
Published: (2024)
by: Lee, Jinwoo, et al.
Published: (2024)
Similar Items
-
Speech Recognition Rescoring with Large Speech-Text Foundation Models
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2024) -
Group Relative Policy Optimization for Speech Recognition
by: Shivakumar, Prashanth Gurunath, et al.
Published: (2025) -
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
by: Pandey, Rahul, et al.
Published: (2023) -
Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
by: Lin, Guan-Ting, et al.
Published: (2024) -
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
by: Everson, Kevin, et al.
Published: (2024)