The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
Fuente:
arXiv
Salvato in:
| Autore principale: | Billa, Jayadev |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
di: Prakash, Jeena, et al.
Pubblicazione: (2025)
di: Prakash, Jeena, et al.
Pubblicazione: (2025)
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
di: Lathouwers, Gus, et al.
Pubblicazione: (2026)
di: Lathouwers, Gus, et al.
Pubblicazione: (2026)
Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
di: Minixhofer, Christoph, et al.
Pubblicazione: (2024)
di: Minixhofer, Christoph, et al.
Pubblicazione: (2024)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
di: Min, Anna, et al.
Pubblicazione: (2025)
di: Min, Anna, et al.
Pubblicazione: (2025)
Articulation-Informed ASR: Integrating Articulatory Features into ASR via Auxiliary Speech Inversion and Cross-Attention Fusion
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2025)
VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)
di: Zhuo, Jianheng, et al.
Pubblicazione: (2025)
Streaming Speech-to-Text Translation with a SpeechLLM
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
di: Parcollet, Titouan, et al.
Pubblicazione: (2026)
Can Speech LLMs Think while Listening?
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
di: Teleki, Maria, et al.
Pubblicazione: (2025)
di: Teleki, Maria, et al.
Pubblicazione: (2025)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024)
di: Seide, Frank, et al.
Pubblicazione: (2024)
Samba-ASR: State-Of-The-Art Speech Recognition Leveraging Structured State-Space Models
di: Shakhadri, Syed Abdul Gaffar, et al.
Pubblicazione: (2025)
di: Shakhadri, Syed Abdul Gaffar, et al.
Pubblicazione: (2025)
Closing the Gap Between Text and Speech Understanding in LLMs
di: Cuervo, Santiago, et al.
Pubblicazione: (2025)
di: Cuervo, Santiago, et al.
Pubblicazione: (2025)
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
di: Srivastav, Vaibhav, et al.
Pubblicazione: (2025)
di: Srivastav, Vaibhav, et al.
Pubblicazione: (2025)
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
di: Ling, Shaoshi, et al.
Pubblicazione: (2025)
di: Ling, Shaoshi, et al.
Pubblicazione: (2025)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2025)
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2025)
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
di: Moure, Pehuén, et al.
Pubblicazione: (2026)
di: Moure, Pehuén, et al.
Pubblicazione: (2026)
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
di: Fan, Xiaoran, et al.
Pubblicazione: (2025)
di: Fan, Xiaoran, et al.
Pubblicazione: (2025)
Enhancing ASR Performance in the Medical Domain for Dravidian Languages
di: Devarakonda, Sri Charan, et al.
Pubblicazione: (2026)
di: Devarakonda, Sri Charan, et al.
Pubblicazione: (2026)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
di: He, Jiajun, et al.
Pubblicazione: (2025)
di: He, Jiajun, et al.
Pubblicazione: (2025)
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
di: Cao, Di, et al.
Pubblicazione: (2026)
di: Cao, Di, et al.
Pubblicazione: (2026)
LoRA-Whisper: Parameter-Efficient and Extensible Multilingual ASR
di: Song, Zheshu, et al.
Pubblicazione: (2024)
di: Song, Zheshu, et al.
Pubblicazione: (2024)
TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
di: Yang, Cheng-Yeh, et al.
Pubblicazione: (2026)
Toward Responsible ASR for African American English Speakers: A Scoping Review of Bias and Equity in Speech Technology
di: Cunningham, Jay L., et al.
Pubblicazione: (2025)
di: Cunningham, Jay L., et al.
Pubblicazione: (2025)
A Comparative Study of LLM-based ASR and Whisper in Low Resource and Code Switching Scenario
di: Song, Zheshu, et al.
Pubblicazione: (2024)
di: Song, Zheshu, et al.
Pubblicazione: (2024)
Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
di: Taherinezhad, Fatemeh, et al.
Pubblicazione: (2025)
di: Taherinezhad, Fatemeh, et al.
Pubblicazione: (2025)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
di: Parulekar, Amruta, et al.
Pubblicazione: (2025)
di: Parulekar, Amruta, et al.
Pubblicazione: (2025)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
Fun-ASR Technical Report
di: An, Keyu, et al.
Pubblicazione: (2025)
di: An, Keyu, et al.
Pubblicazione: (2025)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
di: Fong, Seraphina, et al.
Pubblicazione: (2025)
di: Fong, Seraphina, et al.
Pubblicazione: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024)
di: Jia, Junteng, et al.
Pubblicazione: (2024)
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2025)
di: Yusuyin, Saierdaer, et al.
Pubblicazione: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
Zero-Shot Speech LLMs for Multi-Aspect Evaluation of L2 Speech: Challenges and Opportunities
di: Parikh, Aditya Kamlesh, et al.
Pubblicazione: (2026)
di: Parikh, Aditya Kamlesh, et al.
Pubblicazione: (2026)
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
di: Billa, Jayadev
Pubblicazione: (2026) -
Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
di: Prakash, Jeena, et al.
Pubblicazione: (2025) -
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024) -
Utterance-Level Methods for Identifying Reliable ASR-Output for Child Speech
di: Lathouwers, Gus, et al.
Pubblicazione: (2026) -
Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
di: Minixhofer, Christoph, et al.
Pubblicazione: (2024)