Comparing Discrete and Continuous Space LLMs for Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Yaoxun, Zhang, Shi-Xiong, Yu, Jianwei, Wu, Zhiyong, Yu, Dong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
di: Wang, Dingdong, et al.
Pubblicazione: (2025)
HydraFormer: One Encoder For All Subsampling Rates
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
di: Xu, Yaoxun, et al.
Pubblicazione: (2024)
Exploring Stability-Plasticity Trade-offs for Continual Named Entity Recognition
di: Zhang, Duzhen, et al.
Pubblicazione: (2025)
di: Zhang, Duzhen, et al.
Pubblicazione: (2025)
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
di: Luu, Nam, et al.
Pubblicazione: (2025)
di: Luu, Nam, et al.
Pubblicazione: (2025)
CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Federated Incremental Named Entity Recognition
di: Zhang, Duzhen, et al.
Pubblicazione: (2024)
di: Zhang, Duzhen, et al.
Pubblicazione: (2024)
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
di: Rong, Yiming, et al.
Pubblicazione: (2025)
di: Rong, Yiming, et al.
Pubblicazione: (2025)
Children's Speech Recognition through Discrete Token Enhancement
di: Sukhadia, Vrunda N., et al.
Pubblicazione: (2024)
di: Sukhadia, Vrunda N., et al.
Pubblicazione: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024)
di: Seide, Frank, et al.
Pubblicazione: (2024)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
di: Liu, Heyang, et al.
Pubblicazione: (2025)
di: Liu, Heyang, et al.
Pubblicazione: (2025)
Zero-resource Speech Translation and Recognition with LLMs
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
di: Mundnich, Karel, et al.
Pubblicazione: (2024)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
di: Zhang, Linhao, et al.
Pubblicazione: (2025)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
di: Nassereldine, Amir, et al.
Pubblicazione: (2024)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
LaSR: Context-Aware Speech Recognition via Latent Reasoning
di: Liu, Heyang, et al.
Pubblicazione: (2026)
di: Liu, Heyang, et al.
Pubblicazione: (2026)
Continuously Learning New Words in Automatic Speech Recognition
di: Huber, Christian, et al.
Pubblicazione: (2024)
di: Huber, Christian, et al.
Pubblicazione: (2024)
ReHear: Iterative Pseudo-Label Refinement for Semi-Supervised Speech Recognition via Audio Large Language Models
di: Liu, Zefang, et al.
Pubblicazione: (2026)
di: Liu, Zefang, et al.
Pubblicazione: (2026)
Preference Alignment Improves Language Model-Based TTS
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning
di: Fang, Yangui, et al.
Pubblicazione: (2025)
di: Fang, Yangui, et al.
Pubblicazione: (2025)
Teaching LLMs to Refine with Tools
di: Yu, Dian, et al.
Pubblicazione: (2024)
di: Yu, Dian, et al.
Pubblicazione: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
Continual Adaptation for Pacific Indigenous Speech Recognition
di: Xiao, Yang, et al.
Pubblicazione: (2026)
di: Xiao, Yang, et al.
Pubblicazione: (2026)
Continuous Speech Tokenizer in Text To Speech
di: Li, Yixing, et al.
Pubblicazione: (2024)
di: Li, Yixing, et al.
Pubblicazione: (2024)
Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
di: Tseng, Yuan, et al.
Pubblicazione: (2025)
BANER: Boundary-Aware LLMs for Few-Shot Named Entity Recognition
di: Guo, Quanjiang, et al.
Pubblicazione: (2024)
di: Guo, Quanjiang, et al.
Pubblicazione: (2024)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
di: Xu, Jing, et al.
Pubblicazione: (2026)
di: Xu, Jing, et al.
Pubblicazione: (2026)
Push the Limit of Multi-modal Emotion Recognition by Prompting LLMs with Receptive-Field-Aware Attention Weighting
di: Zhang, Han, et al.
Pubblicazione: (2024)
di: Zhang, Han, et al.
Pubblicazione: (2024)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
di: Yu, Jiazuo, et al.
Pubblicazione: (2024)
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
di: Abu, Turi, et al.
Pubblicazione: (2025)
di: Abu, Turi, et al.
Pubblicazione: (2025)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
di: Wang, Xiaoyang, et al.
Pubblicazione: (2025)
di: Wang, Xiaoyang, et al.
Pubblicazione: (2025)
MM-LLMs: Recent Advances in MultiModal Large Language Models
di: Zhang, Duzhen, et al.
Pubblicazione: (2024)
di: Zhang, Duzhen, et al.
Pubblicazione: (2024)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
di: Park, Keunhyeung, et al.
Pubblicazione: (2025)
di: Park, Keunhyeung, et al.
Pubblicazione: (2025)
A Comparative Study of Continuous Sign Language Recognition Techniques
di: Alyami, Sarah, et al.
Pubblicazione: (2024)
di: Alyami, Sarah, et al.
Pubblicazione: (2024)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
di: Quang, Trung Nguyen, et al.
Pubblicazione: (2026)
di: Quang, Trung Nguyen, et al.
Pubblicazione: (2026)
Augmenting Biomedical Named Entity Recognition with General-domain Resources
di: Yin, Yu, et al.
Pubblicazione: (2024)
di: Yin, Yu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
di: Wang, Dingdong, et al.
Pubblicazione: (2025) -
HydraFormer: One Encoder For All Subsampling Rates
di: Xu, Yaoxun, et al.
Pubblicazione: (2024) -
Exploring Stability-Plasticity Trade-offs for Continual Named Entity Recognition
di: Zhang, Duzhen, et al.
Pubblicazione: (2025) -
End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
di: Luu, Nam, et al.
Pubblicazione: (2025) -
CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
di: Chen, Chen, et al.
Pubblicazione: (2024)