LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ma, Rao, Chen, Tongzhou, Audhkhasi, Kartik, Ramabhadran, Bhuvana |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
STAB: Speech Tokenizer Assessment Benchmark
par: Vashishth, Shikhar, et autres
Publié: (2024)
par: Vashishth, Shikhar, et autres
Publié: (2024)
Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
par: Baskar, Murali Karthick, et autres
Publié: (2024)
par: Baskar, Murali Karthick, et autres
Publié: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
par: Peng, Yifan, et autres
Publié: (2024)
par: Peng, Yifan, et autres
Publié: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
par: Sakuma, Asahi, et autres
Publié: (2025)
par: Sakuma, Asahi, et autres
Publié: (2025)
Unimodal Aggregation for CTC-based Speech Recognition
par: Fang, Ying, et autres
Publié: (2023)
par: Fang, Ying, et autres
Publié: (2023)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
par: Altinok, Duygu
Publié: (2025)
par: Altinok, Duygu
Publié: (2025)
Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
par: Yusuf, Bolaji, et autres
Publié: (2024)
par: Yusuf, Bolaji, et autres
Publié: (2024)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
par: Xie, Yuan, et autres
Publié: (2026)
par: Xie, Yuan, et autres
Publié: (2026)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
par: Wang, Qingzheng, et autres
Publié: (2025)
par: Wang, Qingzheng, et autres
Publié: (2025)
ASTRA: Aligning Speech and Text Representations for Asr without Sampling
par: Gaur, Neeraj, et autres
Publié: (2024)
par: Gaur, Neeraj, et autres
Publié: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
par: Yang, Tzu-Ting, et autres
Publié: (2024)
par: Yang, Tzu-Ting, et autres
Publié: (2024)
The Eloquence team submission for task 1 of MLC-SLM challenge
par: Concina, Lorenzo, et autres
Publié: (2025)
par: Concina, Lorenzo, et autres
Publié: (2025)
Zero-shot Cross-lingual Voice Transfer for TTS
par: Biadsy, Fadi, et autres
Publié: (2024)
par: Biadsy, Fadi, et autres
Publié: (2024)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
par: Sudo, Yui, et autres
Publié: (2024)
par: Sudo, Yui, et autres
Publié: (2024)
Assessment of L2 Oral Proficiency using Speech Large Language Models
par: Ma, Rao, et autres
Publié: (2025)
par: Ma, Rao, et autres
Publié: (2025)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
par: Deng, Keqi, et autres
Publié: (2025)
par: Deng, Keqi, et autres
Publié: (2025)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
par: Zheng, Yuang, et autres
Publié: (2026)
par: Zheng, Yuang, et autres
Publié: (2026)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
par: Liu, Henglyu, et autres
Publié: (2025)
par: Liu, Henglyu, et autres
Publié: (2025)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
par: Le-Duc, Khai, et autres
Publié: (2024)
par: Le-Duc, Khai, et autres
Publié: (2024)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
par: Mdhaffar, Salima, et autres
Publié: (2024)
par: Mdhaffar, Salima, et autres
Publié: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
par: Ma, Rao, et autres
Publié: (2025)
par: Ma, Rao, et autres
Publié: (2025)
Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data
par: Saeki, Takaaki, et autres
Publié: (2024)
par: Saeki, Takaaki, et autres
Publié: (2024)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
par: Fan, Ruchao, et autres
Publié: (2024)
par: Fan, Ruchao, et autres
Publié: (2024)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
par: Raina, Vyas, et autres
Publié: (2024)
par: Raina, Vyas, et autres
Publié: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
par: Huang, Wuwei, et autres
Publié: (2025)
par: Huang, Wuwei, et autres
Publié: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
par: Tsunoo, Emiru, et autres
Publié: (2025)
par: Tsunoo, Emiru, et autres
Publié: (2025)
Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models
par: Ljubešić, Nikola, et autres
Publié: (2025)
par: Ljubešić, Nikola, et autres
Publié: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
par: Kashiwagi, Yosuke, et autres
Publié: (2024)
par: Kashiwagi, Yosuke, et autres
Publié: (2024)
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models
par: Li, Bo, et autres
Publié: (2025)
par: Li, Bo, et autres
Publié: (2025)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
par: Hou, Junfeng, et autres
Publié: (2024)
par: Hou, Junfeng, et autres
Publié: (2024)
RE-LLM: Refining Empathetic Speech-LLM Responses by Integrating Emotion Nuance
par: Chen, Jing-Han, et autres
Publié: (2026)
par: Chen, Jing-Han, et autres
Publié: (2026)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
par: Zhou, Jiaming, et autres
Publié: (2024)
par: Zhou, Jiaming, et autres
Publié: (2024)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
par: Nakagome, Yu, et autres
Publié: (2025)
par: Nakagome, Yu, et autres
Publié: (2025)
ASR Error Correction using Large Language Models
par: Ma, Rao, et autres
Publié: (2024)
par: Ma, Rao, et autres
Publié: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
par: Fang, Qingkai, et autres
Publié: (2024)
par: Fang, Qingkai, et autres
Publié: (2024)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
par: Ao, Junyi, et autres
Publié: (2025)
par: Ao, Junyi, et autres
Publié: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
par: Wang, Yujin, et autres
Publié: (2022)
par: Wang, Yujin, et autres
Publié: (2022)
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
par: Song, Yaodong, et autres
Publié: (2025)
par: Song, Yaodong, et autres
Publié: (2025)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
par: Casanova, Edresson, et autres
Publié: (2024)
par: Casanova, Edresson, et autres
Publié: (2024)
An End-to-End Speech Summarization Using Large Language Model
par: Shang, Hengchao, et autres
Publié: (2024)
par: Shang, Hengchao, et autres
Publié: (2024)
Documents similaires
-
STAB: Speech Tokenizer Assessment Benchmark
par: Vashishth, Shikhar, et autres
Publié: (2024) -
Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
par: Baskar, Murali Karthick, et autres
Publié: (2024) -
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
par: Peng, Yifan, et autres
Publié: (2024) -
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
par: Sakuma, Asahi, et autres
Publié: (2025) -
Unimodal Aggregation for CTC-based Speech Recognition
par: Fang, Ying, et autres
Publié: (2023)