Faster Speech-LLaMA Inference with Multi-token Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raj, Desh, Keren, Gil, Jia, Junteng, Mahadeokar, Jay, Kalinli, Ozlem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
von: Raj, Desh
Veröffentlicht: (2024)
von: Raj, Desh
Veröffentlicht: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
von: Keren, Gil, et al.
Veröffentlicht: (2024)
von: Keren, Gil, et al.
Veröffentlicht: (2024)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)
von: Fang, Qingkai, et al.
Veröffentlicht: (2025)
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
Can Speech LLMs Think while Listening?
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
Fewer-token Neural Speech Codec with Time-invariant Codes
von: Ren, Yong, et al.
Veröffentlicht: (2023)
von: Ren, Yong, et al.
Veröffentlicht: (2023)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2026)
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2026)
On Speaker Attribution with SURT
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
Conversational Speech Naturalness Predictor
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
von: Qin, Chengyuan, et al.
Veröffentlicht: (2025)
von: Qin, Chengyuan, et al.
Veröffentlicht: (2025)
dLLM-ASR: A Faster Diffusion LLM-based Framework for Speech Recognition
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
FreeCodec: A disentangled neural speech codec with fewer tokens
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
von: Bukhari, Hazim, et al.
Veröffentlicht: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
von: Cornell, Samuele, et al.
Veröffentlicht: (2025)
von: Cornell, Samuele, et al.
Veröffentlicht: (2025)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Combined Generative and Predictive Modeling for Speech Super-resolution
von: Wang, Heming, et al.
Veröffentlicht: (2024)
von: Wang, Heming, et al.
Veröffentlicht: (2024)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2026)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
von: Seide, Frank, et al.
Veröffentlicht: (2024)
von: Seide, Frank, et al.
Veröffentlicht: (2024)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
Multi-modal Speech Enhancement with Limited Electromyography Channels
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025)
von: Feng, Fuyuan, et al.
Veröffentlicht: (2025)
Unsupervised Multi-channel Speech Dereverberation via Diffusion
von: Wu, Yulun, et al.
Veröffentlicht: (2025)
von: Wu, Yulun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024) -
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024) -
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024) -
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024) -
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
von: Raj, Desh
Veröffentlicht: (2024)