Speech LLMs are Contextual Reasoning Transcribers
Fuente:
arXiv
Salvato in:
| Autori principali: | Deng, Keqi, Fan, Ruchao, Ren, Bo, Wang, Yiming, Li, Jinyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
di: Deng, Keqi, et al.
Pubblicazione: (2026)
di: Deng, Keqi, et al.
Pubblicazione: (2026)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
di: Wu, Haibin, et al.
Pubblicazione: (2025)
di: Wu, Haibin, et al.
Pubblicazione: (2025)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
di: Ren, Bo, et al.
Pubblicazione: (2026)
di: Ren, Bo, et al.
Pubblicazione: (2026)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2025)
di: Deng, Keqi, et al.
Pubblicazione: (2025)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
di: Zhao, Rui, et al.
Pubblicazione: (2024)
di: Zhao, Rui, et al.
Pubblicazione: (2024)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
di: Ren, Bo, et al.
Pubblicazione: (2025)
di: Ren, Bo, et al.
Pubblicazione: (2025)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
di: Li, Chak-Fai, et al.
Pubblicazione: (2024)
di: Li, Chak-Fai, et al.
Pubblicazione: (2024)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
di: Du, Yexing, et al.
Pubblicazione: (2024)
di: Du, Yexing, et al.
Pubblicazione: (2024)
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
di: Rong, Yiming, et al.
Pubblicazione: (2025)
di: Rong, Yiming, et al.
Pubblicazione: (2025)
TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech
di: Venturini, Shamira, et al.
Pubblicazione: (2026)
di: Venturini, Shamira, et al.
Pubblicazione: (2026)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
di: Moritz, Niko, et al.
Pubblicazione: (2024)
di: Moritz, Niko, et al.
Pubblicazione: (2024)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
di: Wang, Qiongqiong, et al.
Pubblicazione: (2025)
Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer Enhancement
di: Yuan, Xiaowei, et al.
Pubblicazione: (2025)
di: Yuan, Xiaowei, et al.
Pubblicazione: (2025)
Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal
di: Gauthier, Elodie, et al.
Pubblicazione: (2024)
di: Gauthier, Elodie, et al.
Pubblicazione: (2024)
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
di: Du, Yexing, et al.
Pubblicazione: (2025)
di: Du, Yexing, et al.
Pubblicazione: (2025)
OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics
di: Chu, Wei, et al.
Pubblicazione: (2025)
di: Chu, Wei, et al.
Pubblicazione: (2025)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
di: Wang, Yiming, et al.
Pubblicazione: (2023)
di: Wang, Yiming, et al.
Pubblicazione: (2023)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
di: Chen, Guoguo, et al.
Pubblicazione: (2021)
di: Chen, Guoguo, et al.
Pubblicazione: (2021)
From Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMs
di: Guo, Xiaoyong, et al.
Pubblicazione: (2026)
di: Guo, Xiaoyong, et al.
Pubblicazione: (2026)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
DiDOTS: Knowledge Distillation from Large-Language-Models for Dementia Obfuscation in Transcribed Speech
di: Woszczyk, Dominika, et al.
Pubblicazione: (2024)
di: Woszczyk, Dominika, et al.
Pubblicazione: (2024)
Advancing Speech Summarization in Multi-modal LLMs with Reinforcement Learning
di: Ling, Shaoshi, et al.
Pubblicazione: (2025)
di: Ling, Shaoshi, et al.
Pubblicazione: (2025)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
di: Li, Song, et al.
Pubblicazione: (2024)
di: Li, Song, et al.
Pubblicazione: (2024)
Rubato: Transcribing Piano Music with Timestamps
di: Tamer, Nazif Can, et al.
Pubblicazione: (2026)
di: Tamer, Nazif Can, et al.
Pubblicazione: (2026)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
di: Sun, Siqi, et al.
Pubblicazione: (2024)
di: Sun, Siqi, et al.
Pubblicazione: (2024)
ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure
di: Deng, Jie, et al.
Pubblicazione: (2026)
di: Deng, Jie, et al.
Pubblicazione: (2026)
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
di: Chen, Jingyi, et al.
Pubblicazione: (2025)
di: Chen, Jingyi, et al.
Pubblicazione: (2025)
SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
di: Hacioglu, Kadri, et al.
Pubblicazione: (2025)
di: Hacioglu, Kadri, et al.
Pubblicazione: (2025)
Closing the Modality Reasoning Gap for Speech Large Language Models
di: Wang, Chaoren, et al.
Pubblicazione: (2026)
di: Wang, Chaoren, et al.
Pubblicazione: (2026)
Slot Filling as a Reasoning Task for SpeechLLMs
di: Hacioglu, Kadri, et al.
Pubblicazione: (2025)
di: Hacioglu, Kadri, et al.
Pubblicazione: (2025)
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
di: Wang, Siyuan, et al.
Pubblicazione: (2024)
di: Wang, Siyuan, et al.
Pubblicazione: (2024)
Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
di: Ògúnrèmí, Tolúlopé, et al.
Pubblicazione: (2025)
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
di: Shim, Ryan Soh-Eun, et al.
Pubblicazione: (2024)
di: Shim, Ryan Soh-Eun, et al.
Pubblicazione: (2024)
Towards Event Extraction from Speech with Contextual Clues
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
di: Kang, Jingqi, et al.
Pubblicazione: (2024)
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2024)
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2024)
Documenti analoghi
-
UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
di: Deng, Keqi, et al.
Pubblicazione: (2026) -
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
di: Wu, Haibin, et al.
Pubblicazione: (2025) -
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
di: Ren, Bo, et al.
Pubblicazione: (2026) -
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2025) -
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
di: Zhao, Rui, et al.
Pubblicazione: (2024)