Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Mutian, Garner, Philip N.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929758076403712
author He, Mutian
Garner, Philip N.
author_facet He, Mutian
Garner, Philip N.
contents Architectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often unavailable, especially in non-text domains. To remedy this, we present a Cross-Architecture Layerwise Distillation (CALD) approach that jointly converts a transformer model to a linear time substitute and fine-tunes it to a target task. We also compare several means to guide the fine-tuning to optimally retain the desired inference capability from the original model. The methods differ in their use of the target model and the trajectory of the parameters. In a series of empirical studies on language processing, language modeling, and speech processing, we show that CALD can effectively recover the result of the original model, and that the guiding strategy contributes to the result. Some reasons for the variation are suggested.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06846
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
He, Mutian
Garner, Philip N.
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Architectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often unavailable, especially in non-text domains. To remedy this, we present a Cross-Architecture Layerwise Distillation (CALD) approach that jointly converts a transformer model to a linear time substitute and fine-tunes it to a target task. We also compare several means to guide the fine-tuning to optimally retain the desired inference capability from the original model. The methods differ in their use of the target model and the trajectory of the parameters. In a series of empirical studies on language processing, language modeling, and speech processing, we show that CALD can effectively recover the result of the original model, and that the guiding strategy contributes to the result. Some reasons for the variation are suggested.
title Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.06846