Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Qingran, Zhao, Botao, Kang, Zuheng, Li, Xue, He, Yayun, Liu, Chuhang, Zhang, Xulong, Qu, Xiaoyang, Peng, Junqing, Wang, Jianzong
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911416050515968
author Yang, Qingran
Zhao, Botao
Kang, Zuheng
Li, Xue
He, Yayun
Liu, Chuhang
Zhang, Xulong
Qu, Xiaoyang
Peng, Junqing
Wang, Jianzong
author_facet Yang, Qingran
Zhao, Botao
Kang, Zuheng
Li, Xue
He, Yayun
Liu, Chuhang
Zhang, Xulong
Qu, Xiaoyang
Peng, Junqing
Wang, Jianzong
contents The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Knowledge Distillation is effective for LALM compression, existing methods remain underexplored in distilling the cross-modal projection module (Projector), and often struggle with alignment due to differences in feature dimensions. We propose PL-Distill, a KD framework that combines Projector-Level Distillation (PDist) to align audio embeddings and Logits-Level Distillation (LDist) to align output logits. PDist introduces Attention-weighted Centered Kernel Alignment, a novel approach we propose to highlight important time steps and address dimension mismatches. Meanwhile, LDist minimizes the Kullback-Leibler divergence between teacher and student logits from audio and text modalities. On IEMOCAP, RAVDESS, and SAVEE, PL-Distill compresses an 8.4B-parameter teacher to a compact 1.1B-parameter student, consistently outperforming the teacher, state-of-the-art pretrained models, and other KD baselines across all metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01547
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
Yang, Qingran
Zhao, Botao
Kang, Zuheng
Li, Xue
He, Yayun
Liu, Chuhang
Zhang, Xulong
Qu, Xiaoyang
Peng, Junqing
Wang, Jianzong
Sound
Audio and Speech Processing
The emergence of Large Audio-Language Models (LALMs) has advanced Speech Emotion Recognition (SER), but their size limits deployment in resource-constrained environments. While Knowledge Distillation is effective for LALM compression, existing methods remain underexplored in distilling the cross-modal projection module (Projector), and often struggle with alignment due to differences in feature dimensions. We propose PL-Distill, a KD framework that combines Projector-Level Distillation (PDist) to align audio embeddings and Logits-Level Distillation (LDist) to align output logits. PDist introduces Attention-weighted Centered Kernel Alignment, a novel approach we propose to highlight important time steps and address dimension mismatches. Meanwhile, LDist minimizes the Kullback-Leibler divergence between teacher and student logits from audio and text modalities. On IEMOCAP, RAVDESS, and SAVEE, PL-Distill compresses an 8.4B-parameter teacher to a compact 1.1B-parameter student, consistently outperforming the teacher, state-of-the-art pretrained models, and other KD baselines across all metrics.
title Attention-weighted Centered Kernel Alignment for Knowledge Distillation in Large Audio-Language Models Applied to Speech Emotion Recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2602.01547