Gespeichert in:
| Hauptverfasser: | Shih, Yi-Jen, Raj, Desh, Wu, Chunyang, Zhou, Wei, Bong, SK, Gaur, Yashesh, Mahadeokar, Jay, Kalinli, Ozlem, Seltzer, Mike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.07497 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
von: Kang, Wonjune, et al.
Veröffentlicht: (2024)
Conversational Speech Naturalness Predictor
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)
Effective internal language model training and fusion for factorized transducer model
von: Guo, Jinxi, et al.
Veröffentlicht: (2024)
von: Guo, Jinxi, et al.
Veröffentlicht: (2024)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
von: Raj, Desh
Veröffentlicht: (2024)
von: Raj, Desh
Veröffentlicht: (2024)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Efficient Streaming LLM for Speech Recognition
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
von: Jia, Junteng, et al.
Veröffentlicht: (2024)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2026)
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2026)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
von: Seide, Frank, et al.
Veröffentlicht: (2024)
von: Seide, Frank, et al.
Veröffentlicht: (2024)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
von: Xie, Jiamin, et al.
Veröffentlicht: (2023)
Towards measuring fairness in speech recognition: Fair-Speech dataset
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
von: Veliche, Irina-Elena, et al.
Veröffentlicht: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
von: Keren, Gil, et al.
Veröffentlicht: (2024)
von: Keren, Gil, et al.
Veröffentlicht: (2024)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
von: Ma, Yingyi, et al.
Veröffentlicht: (2024)
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
von: Moritz, Niko, et al.
Veröffentlicht: (2024)
von: Moritz, Niko, et al.
Veröffentlicht: (2024)
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
von: Pan, Jing, et al.
Veröffentlicht: (2023)
von: Pan, Jing, et al.
Veröffentlicht: (2023)
Unifying Model and Layer Fusion for Speech Foundation Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
Listen, Think, and Understand
von: Gong, Yuan, et al.
Veröffentlicht: (2023)
von: Gong, Yuan, et al.
Veröffentlicht: (2023)
Latent Speech-Text Transformer
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
von: Yang, Hsiang-Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Hsiang-Cheng, et al.
Veröffentlicht: (2026)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
On Speaker Attribution with SURT
von: Raj, Desh, et al.
Veröffentlicht: (2024)
von: Raj, Desh, et al.
Veröffentlicht: (2024)
DIFFA: Large Language Diffusion Models Can Listen and Understand
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
von: Jeong, Jihoon, et al.
Veröffentlicht: (2026)
Evaluating Speech Enhancement Systems Through Listening Effort
von: Gelderblom, Femke B., et al.
Veröffentlicht: (2024)
von: Gelderblom, Femke B., et al.
Veröffentlicht: (2024)
Towards scalable efficient on-device ASR with transfer learning
von: Pandey, Laxmi, et al.
Veröffentlicht: (2024)
von: Pandey, Laxmi, et al.
Veröffentlicht: (2024)
French Listening Tests for the Assessment of Intelligibility, Quality, and Identity of Body-Conducted Speech Enhancement
von: Joubaud, Thomas, et al.
Veröffentlicht: (2025)
von: Joubaud, Thomas, et al.
Veröffentlicht: (2025)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
von: Zhang, Lin, et al.
Veröffentlicht: (2026)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
von: Hu, Cheng-Hung, et al.
Veröffentlicht: (2025)
Can Masked Autoencoders Also Listen to Birds?
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
von: Cornell, Samuele, et al.
Veröffentlicht: (2025)
von: Cornell, Samuele, et al.
Veröffentlicht: (2025)
Enhanced Generative Machine Listener
von: Raj, Vishnu, et al.
Veröffentlicht: (2025)
von: Raj, Vishnu, et al.
Veröffentlicht: (2025)
Evaluation of an ITD-to-ILD Transformation as a Method to Restore the Spatial Benefit in Speech Intelligibility in Hearing Impaired Listeners
von: Bäumer, Timm-Jonas, et al.
Veröffentlicht: (2025)
von: Bäumer, Timm-Jonas, et al.
Veröffentlicht: (2025)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
Meta-PerSER: Few-Shot Listener Personalized Speech Emotion Recognition via Meta-learning
von: Shen, Liang-Yeh, et al.
Veröffentlicht: (2025)
von: Shen, Liang-Yeh, et al.
Veröffentlicht: (2025)
Listening Between the Lines: Synthetic Speech Detection Disregarding Verbal Content
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
von: Salvi, Davide, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Faster Speech-LLaMA Inference with Multi-token Prediction
von: Raj, Desh, et al.
Veröffentlicht: (2024) -
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
von: Kang, Wonjune, et al.
Veröffentlicht: (2024) -
Conversational Speech Naturalness Predictor
von: Xu, Anfeng, et al.
Veröffentlicht: (2026) -
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024) -
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
von: Yang, Yufeng, et al.
Veröffentlicht: (2024)