Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mu, Bingshen, Liu, Hexin, Xue, Hongfei, Wei, Kun, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
von: Mu, Bingshen, et al.
Veröffentlicht: (2026)
von: Mu, Bingshen, et al.
Veröffentlicht: (2026)
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
von: Geng, Xuelong, et al.
Veröffentlicht: (2024)
von: Geng, Xuelong, et al.
Veröffentlicht: (2024)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
CLAR: CIF-Localized Alignment for Retrieval-Augmented Speech LLM-Based Contextual ASR
von: Huang, Shangkun, et al.
Veröffentlicht: (2026)
von: Huang, Shangkun, et al.
Veröffentlicht: (2026)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2026)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
von: Kamble, Anand, et al.
Veröffentlicht: (2023)
von: Kamble, Anand, et al.
Veröffentlicht: (2023)
Ideal-LLM: Integrating Dual Encoders and Language-Adapted LLM for Multilingual Speech-to-Text
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
CoHear: Conversation Enhancement via Multi-Earphone Collaboration
von: He, Lixing, et al.
Veröffentlicht: (2025)
von: He, Lixing, et al.
Veröffentlicht: (2025)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
ASR Benchmarking: Need for a More Representative Conversational Dataset
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
von: Xie, Yuan, et al.
Veröffentlicht: (2026)
OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
von: Liao, Yujie, et al.
Veröffentlicht: (2026)
von: Liao, Yujie, et al.
Veröffentlicht: (2026)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
von: Kirton-Wingate, Jasper, et al.
Veröffentlicht: (2024)
Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
von: Shi, Hao, et al.
Veröffentlicht: (2026)
von: Shi, Hao, et al.
Veröffentlicht: (2026)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech Test for Diagnosing Presbycusis
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
Less is More: Data Curation Matters in Scaling Speech Enhancement
von: Li, Chenda, et al.
Veröffentlicht: (2025)
von: Li, Chenda, et al.
Veröffentlicht: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Identifying Hearing Difficulty Moments in Conversational Audio
von: Collins, Jack, et al.
Veröffentlicht: (2025)
von: Collins, Jack, et al.
Veröffentlicht: (2025)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2026)
The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2026)
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025) -
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025) -
HDMoLE: Mixture of LoRA Experts with Hierarchical Routing and Dynamic Thresholds for Fine-Tuning LLM-based ASR Models
von: Mu, Bingshen, et al.
Veröffentlicht: (2024) -
LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
von: Mu, Bingshen, et al.
Veröffentlicht: (2026) -
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
von: Geng, Xuelong, et al.
Veröffentlicht: (2024)