Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ziwei, Dong, Lukuang, Yusuyin, Saierdaer, Zhao, Xianyu, Ou, Zhijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2025)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2025)
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
von: Zhao, Wenbo, et al.
Veröffentlicht: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025)
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2025)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
von: Wu, Ya-Tse, et al.
Veröffentlicht: (2026)
von: Wu, Ya-Tse, et al.
Veröffentlicht: (2026)
Energy-Based Models with Applications to Speech and Language Processing
von: Ou, Zhijian
Veröffentlicht: (2024)
von: Ou, Zhijian
Veröffentlicht: (2024)
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2024)
von: Kong, Xiangzhu, et al.
Veröffentlicht: (2024)
Efficient Scaling for LLM-based ASR
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
A Bottom-up Framework with Language-universal Speech Attribute Modeling for Syllable-based ASR
von: Yen, Hao, et al.
Veröffentlicht: (2025)
von: Yen, Hao, et al.
Veröffentlicht: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2024)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2024)
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
von: Salvi, Davide, et al.
Veröffentlicht: (2025)
von: Salvi, Davide, et al.
Veröffentlicht: (2025)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
von: Novitasari, Sashi, et al.
Veröffentlicht: (2026)
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
von: Tao, Dehua, et al.
Veröffentlicht: (2026)
von: Tao, Dehua, et al.
Veröffentlicht: (2026)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
von: Li, Li, et al.
Veröffentlicht: (2026)
von: Li, Li, et al.
Veröffentlicht: (2026)
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
von: Monir, Nasser-Eddine, et al.
Veröffentlicht: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026)
von: Saon, George, et al.
Veröffentlicht: (2026)
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
Towards a Single ASR Model That Generalizes to Disordered Speech
von: Tobin, Jimmy, et al.
Veröffentlicht: (2024)
von: Tobin, Jimmy, et al.
Veröffentlicht: (2024)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
von: Pulikodan, Sujith, et al.
Veröffentlicht: (2025)
von: Pulikodan, Sujith, et al.
Veröffentlicht: (2025)
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
von: Wang, Tianzi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026) -
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2025) -
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
von: Liu, Hanwen, et al.
Veröffentlicht: (2026) -
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025) -
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)