Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Eungbeom, Kim, Hantae, Lee, Kyogu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2406.07909
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917691584937984
author Kim, Eungbeom
Kim, Hantae
Lee, Kyogu
author_facet Kim, Eungbeom
Kim, Hantae
Lee, Kyogu
contents Transformer encoder with connectionist temporal classification (CTC) framework is widely used for automatic speech recognition (ASR). However, knowledge distillation (KD) for ASR displays a problem of disagreement between teacher-student models in frame-level alignment which ultimately hinders it from improving the student model's performance. In order to resolve this problem, this paper introduces a self-knowledge distillation (SKD) method that guides the frame-level alignment during the training time. In contrast to the conventional method using separate teacher and student models, this study introduces a simple and effective method sharing encoder layers and applying the sub-model as the student model. Overall, our approach is effective in improving both the resource efficiency as well as performance. We also conducted an experimental analysis of the spike timings to illustrate that the proposed method improves performance by reducing the alignment disagreement.
format Preprint
id arxiv_https___arxiv_org_abs_2406_07909
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
Kim, Eungbeom
Kim, Hantae
Lee, Kyogu
Audio and Speech Processing
Computation and Language
Sound
Machine Learning
Transformer encoder with connectionist temporal classification (CTC) framework is widely used for automatic speech recognition (ASR). However, knowledge distillation (KD) for ASR displays a problem of disagreement between teacher-student models in frame-level alignment which ultimately hinders it from improving the student model's performance. In order to resolve this problem, this paper introduces a self-knowledge distillation (SKD) method that guides the frame-level alignment during the training time. In contrast to the conventional method using separate teacher and student models, this study introduces a simple and effective method sharing encoder layers and applying the sub-model as the student model. Overall, our approach is effective in improving both the resource efficiency as well as performance. We also conducted an experimental analysis of the spike timings to illustrate that the proposed method improves performance by reducing the alignment disagreement.
title Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
topic Audio and Speech Processing
Computation and Language
Sound
Machine Learning
url https://arxiv.org/abs/2406.07909