X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Cao, Di, Fu, Dongjie, Yu, Hai, Zheng, Siqi, Tan, Xu, Jin, Tao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
por: Liu, Wei, et al.
Publicado: (2025)
por: Liu, Wei, et al.
Publicado: (2025)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
por: Ueda, Lucas, et al.
Publicado: (2025)
por: Ueda, Lucas, et al.
Publicado: (2025)
TASU: Text-Only Alignment for Speech Understanding
por: Peng, Jing, et al.
Publicado: (2025)
por: Peng, Jing, et al.
Publicado: (2025)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025)
por: Yuan, Xihao, et al.
Publicado: (2025)
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
por: Fu, Li, et al.
Publicado: (2024)
por: Fu, Li, et al.
Publicado: (2024)
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
por: Ok, Seaone, et al.
Publicado: (2026)
por: Ok, Seaone, et al.
Publicado: (2026)
SSR: Alignment-Aware Modality Connector for Speech Language Models
por: Tan, Weiting, et al.
Publicado: (2024)
por: Tan, Weiting, et al.
Publicado: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
por: Fan, Cunhang, et al.
Publicado: (2023)
por: Fan, Cunhang, et al.
Publicado: (2023)
Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation
por: Wang, Zheng, et al.
Publicado: (2026)
por: Wang, Zheng, et al.
Publicado: (2026)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
por: Tao, Dehua, et al.
Publicado: (2024)
por: Tao, Dehua, et al.
Publicado: (2024)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
por: Choi, Jeongsoo, et al.
Publicado: (2025)
por: Choi, Jeongsoo, et al.
Publicado: (2025)
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
por: Tao, Dehua, et al.
Publicado: (2026)
por: Tao, Dehua, et al.
Publicado: (2026)
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
por: Liu, Dong, et al.
Publicado: (2025)
por: Liu, Dong, et al.
Publicado: (2025)
Complex Recurrent Variational Autoencoder with Application to Speech Enhancement
por: Xie, Yuying, et al.
Publicado: (2022)
por: Xie, Yuying, et al.
Publicado: (2022)
ARTT: Augmented Reverberant-Target Training for Unsupervised Monaural Speech Dereverberation
por: Song, Siqi, et al.
Publicado: (2026)
por: Song, Siqi, et al.
Publicado: (2026)
Multi-Distillation from Speech and Music Representation Models
por: Wei, Jui-Chiang, et al.
Publicado: (2025)
por: Wei, Jui-Chiang, et al.
Publicado: (2025)
Exploring the Capability of Mamba in Speech Applications
por: Miyazaki, Koichi, et al.
Publicado: (2024)
por: Miyazaki, Koichi, et al.
Publicado: (2024)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
por: Yuan, Ze, et al.
Publicado: (2024)
por: Yuan, Ze, et al.
Publicado: (2024)
TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMs
por: Peng, Jing, et al.
Publicado: (2026)
por: Peng, Jing, et al.
Publicado: (2026)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
por: Liu, Wenrui, et al.
Publicado: (2025)
por: Liu, Wenrui, et al.
Publicado: (2025)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
por: Ma, Linhan, et al.
Publicado: (2025)
por: Ma, Linhan, et al.
Publicado: (2025)
LLMs and Speech: Integration vs. Combination
por: Schmitt, Robin, et al.
Publicado: (2026)
por: Schmitt, Robin, et al.
Publicado: (2026)
Robust One-step Speech Enhancement via Consistency Distillation
por: Xu, Liang, et al.
Publicado: (2025)
por: Xu, Liang, et al.
Publicado: (2025)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
por: Zhang, Hanlin, et al.
Publicado: (2026)
por: Zhang, Hanlin, et al.
Publicado: (2026)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
por: Han, Runduo, et al.
Publicado: (2024)
por: Han, Runduo, et al.
Publicado: (2024)
Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
por: Yang, Qingran, et al.
Publicado: (2026)
por: Yang, Qingran, et al.
Publicado: (2026)
Group Relative Policy Optimization for Speech Recognition
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2025)
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2025)
AS-Speech: Adaptive Style For Speech Synthesis
por: Li, Zhipeng, et al.
Publicado: (2024)
por: Li, Zhipeng, et al.
Publicado: (2024)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
por: Cui, Mingyu, et al.
Publicado: (2025)
por: Cui, Mingyu, et al.
Publicado: (2025)
Text-aware Speech Separation for Multi-talker Keyword Spotting
por: Li, Haoyu, et al.
Publicado: (2024)
por: Li, Haoyu, et al.
Publicado: (2024)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
por: Lu, Ke-Han, et al.
Publicado: (2024)
por: Lu, Ke-Han, et al.
Publicado: (2024)
Adaptive Duration Model for Text Speech Alignment
por: Cao, Junjie
Publicado: (2025)
por: Cao, Junjie
Publicado: (2025)
Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models
por: C, Shiva Kumar, et al.
Publicado: (2025)
por: C, Shiva Kumar, et al.
Publicado: (2025)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
por: Niu, Zhikang, et al.
Publicado: (2025)
por: Niu, Zhikang, et al.
Publicado: (2025)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
por: Eom, SooHwan, et al.
Publicado: (2026)
por: Eom, SooHwan, et al.
Publicado: (2026)
Validating Computational Markers of Depressive Behavior: Cross-Linguistic Speech-Based Depression Detection with Neurophysiological Validation
por: Tao, Fuxiang, et al.
Publicado: (2026)
por: Tao, Fuxiang, et al.
Publicado: (2026)
AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
por: Kalkhorani, Vahid Ahmadi, et al.
Publicado: (2024)
por: Kalkhorani, Vahid Ahmadi, et al.
Publicado: (2024)
Audio-Image Cross-Modal Retrieval with Onomatopoeic Images
por: Imoto, Keisuke, et al.
Publicado: (2026)
por: Imoto, Keisuke, et al.
Publicado: (2026)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
por: Hsu, Ming-Hao, et al.
Publicado: (2026)
por: Hsu, Ming-Hao, et al.
Publicado: (2026)
DISPATCH: Distilling Selective Patches for Speech Enhancement
por: Kim, Dohwan, et al.
Publicado: (2025)
por: Kim, Dohwan, et al.
Publicado: (2025)
Ejemplares similares
-
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
por: Liu, Wei, et al.
Publicado: (2025) -
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
por: Ueda, Lucas, et al.
Publicado: (2025) -
TASU: Text-Only Alignment for Speech Understanding
por: Peng, Jing, et al.
Publicado: (2025) -
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025) -
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition
por: Fu, Li, et al.
Publicado: (2024)