Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gopal, Shreyas, Wu, Donghang, Anshul, Ashutosh, Heng, Yeo Yue, Peng, Yizhou, Li, Haoyang, Liu, Hexin, Chng, Eng Siong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
von: Gopal, Shreyas, et al.
Veröffentlicht: (2025)
von: Gopal, Shreyas, et al.
Veröffentlicht: (2025)
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Improving Code-Switching Speech Recognition with TTS Data Augmentation
von: Yeo, Yue Heng, et al.
Veröffentlicht: (2026)
von: Yeo, Yue Heng, et al.
Veröffentlicht: (2026)
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
von: Anshul, Ashutosh, et al.
Veröffentlicht: (2025)
von: Anshul, Ashutosh, et al.
Veröffentlicht: (2025)
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
NTU Speechlab LLM-Based Multilingual ASR System for Interspeech MLC-SLM Challenge 2025
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2024)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
von: Liu, Changsong, et al.
Veröffentlicht: (2025)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
Code-switching Speech Recognition Under the Lens: Model- and Data-Centric Perspectives
von: Liu, Hexin, et al.
Veröffentlicht: (2025)
von: Liu, Hexin, et al.
Veröffentlicht: (2025)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Aligning Speech to Languages to Enhance Code-switching Speech Recognition
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
von: Liu, Hexin, et al.
Veröffentlicht: (2024)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
Prosodic Boundary-Aware Streaming Generation for LLM-Based TTS with Streaming Text Input
von: Liu, Changsong, et al.
Veröffentlicht: (2026)
von: Liu, Changsong, et al.
Veröffentlicht: (2026)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
von: Zhang, Haoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2026)
StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
von: Luong, Hieu-Thi, et al.
Veröffentlicht: (2024)
Noise-Aware Speech Separation with Contrastive Learning
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection
von: Li, Yuxin, et al.
Veröffentlicht: (2026)
von: Li, Yuxin, et al.
Veröffentlicht: (2026)
Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Text-based Talking Video Editing with Cascaded Conditional Diffusion
von: Han, Bo, et al.
Veröffentlicht: (2024)
von: Han, Bo, et al.
Veröffentlicht: (2024)
Improving Synthetic Data Training for Contextual Biasing Models with a Keyword-Aware Cost Function
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2025)
von: Kwok, Chin Yuen, et al.
Veröffentlicht: (2025)
Noise-aware Speech Enhancement using Diffusion Probabilistic Model
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
Speech Separation using Neural Audio Codecs with Embedding Loss
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Towards Audio Codec-based Speech Separation
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
von: Gopal, Shreyas, et al.
Veröffentlicht: (2025) -
Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR
von: Peng, Yizhou, et al.
Veröffentlicht: (2025) -
Improving Code-Switching Speech Recognition with TTS Data Augmentation
von: Yeo, Yue Heng, et al.
Veröffentlicht: (2026) -
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
von: Anshul, Ashutosh, et al.
Veröffentlicht: (2025) -
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)