Gespeichert in:
| Hauptverfasser: | Guo, Jinxi, Moritz, Niko, Ma, Yingyi, Seide, Frank, Wu, Chunyang, Mahadeokar, Jay, Kalinli, Ozlem, Fuegen, Christian, Seltzer, Mike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.01716 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
Can Speech LLMs Think while Listening?
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
von: Lin, Ju, et al.
Veröffentlicht: (2024)
von: Lin, Ju, et al.
Veröffentlicht: (2024)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
von: Moritz, Niko, et al.
Veröffentlicht: (2024)
von: Moritz, Niko, et al.
Veröffentlicht: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2026)
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2026)
Towards scalable efficient on-device ASR with transfer learning
von: Pandey, Laxmi, et al.
Veröffentlicht: (2024)
von: Pandey, Laxmi, et al.
Veröffentlicht: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
von: Keren, Gil, et al.
Veröffentlicht: (2024)
von: Keren, Gil, et al.
Veröffentlicht: (2024)
Towards measuring fairness in speech recognition: Fair-Speech dataset
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
von: Seide, Frank, et al.
Veröffentlicht: (2024)
von: Seide, Frank, et al.
Veröffentlicht: (2024)
Conversational Speech Naturalness Predictor
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024)
von: Rabatin, Rastislav, et al.
Veröffentlicht: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
von: Shen, Maohao, et al.
Veröffentlicht: (2024)
von: Shen, Maohao, et al.
Veröffentlicht: (2024)
Towards audio language modeling -- an overview
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
von: Samui, Suman
Veröffentlicht: (2022)
von: Samui, Suman
Veröffentlicht: (2022)
AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs
von: Fathullah, Yassir, et al.
Veröffentlicht: (2023)
von: Fathullah, Yassir, et al.
Veröffentlicht: (2023)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
von: Pandey, Rahul, et al.
Veröffentlicht: (2023)
Progressive unsupervised domain adaptation for ASR using ensemble models and multi-stage training
von: Ahmad, Rehan, et al.
Veröffentlicht: (2024)
von: Ahmad, Rehan, et al.
Veröffentlicht: (2024)
Selecting N-lowest scores for training MOS prediction models
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
von: Kondo, Yuto, et al.
Veröffentlicht: (2025)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
Building English ASR model with regional language support
von: Agrawal, Purvi, et al.
Veröffentlicht: (2025)
von: Agrawal, Purvi, et al.
Veröffentlicht: (2025)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Encoding of lexical tone in self-supervised models of spoken language
von: Shen, Gaofei, et al.
Veröffentlicht: (2024)
von: Shen, Gaofei, et al.
Veröffentlicht: (2024)
How to train your ears: Auditory-model emulation for large-dynamic-range inputs and mild-to-severe hearing losses
von: Leer, Peter, et al.
Veröffentlicht: (2024)
von: Leer, Peter, et al.
Veröffentlicht: (2024)
Mellow: a small audio language model for reasoning
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
von: Li, Yiming, et al.
Veröffentlicht: (2024)
von: Li, Yiming, et al.
Veröffentlicht: (2024)
A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data
von: Tran, Minh, et al.
Veröffentlicht: (2025)
von: Tran, Minh, et al.
Veröffentlicht: (2025)
Word-wise intonation model for cross-language TTS systems
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
von: A., Tomilov A., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
von: Deng, Keqi, et al.
Veröffentlicht: (2024) -
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024) -
Can Speech LLMs Think while Listening?
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025) -
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024) -
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)