PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Tzu-Han, Wang, How-Shing, Weng, Hao-Yung, Peng, Kuang-Chen, Chen, Zih-Ching, Lee, Hung-yi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913225469067264
author Lin, Tzu-Han
Wang, How-Shing
Weng, Hao-Yung
Peng, Kuang-Chen
Chen, Zih-Ching
Lee, Hung-yi
author_facet Lin, Tzu-Han
Wang, How-Shing
Weng, Hao-Yung
Peng, Kuang-Chen
Chen, Zih-Ching
Lee, Hung-yi
contents Parameter-Efficient Fine-Tuning (PEFT) is increasingly recognized as an effective method in speech processing. However, the optimal approach and the placement of PEFT methods remain inconclusive. Our study conducts extensive experiments to compare different PEFT methods and their layer-wise placement adapting Differentiable Architecture Search (DARTS). We also explore the use of ensemble learning to leverage diverse PEFT strategies. The results reveal that DARTS does not outperform the baseline approach, which involves inserting the same PEFT method into all layers of a Self-Supervised Learning (SSL) model. In contrast, an ensemble learning approach, particularly one employing majority voting, demonstrates superior performance. Our statistical evidence indicates that different PEFT methods learn in varied ways. This variation might explain why the synergistic integration of various PEFT methods through ensemble learning can harness their unique learning capabilities more effectively compared to individual layer-wise optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02122
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
Lin, Tzu-Han
Wang, How-Shing
Weng, Hao-Yung
Peng, Kuang-Chen
Chen, Zih-Ching
Lee, Hung-yi
Computation and Language
Sound
Audio and Speech Processing
Parameter-Efficient Fine-Tuning (PEFT) is increasingly recognized as an effective method in speech processing. However, the optimal approach and the placement of PEFT methods remain inconclusive. Our study conducts extensive experiments to compare different PEFT methods and their layer-wise placement adapting Differentiable Architecture Search (DARTS). We also explore the use of ensemble learning to leverage diverse PEFT strategies. The results reveal that DARTS does not outperform the baseline approach, which involves inserting the same PEFT method into all layers of a Self-Supervised Learning (SSL) model. In contrast, an ensemble learning approach, particularly one employing majority voting, demonstrates superior performance. Our statistical evidence indicates that different PEFT methods learn in varied ways. This variation might explain why the synergistic integration of various PEFT methods through ensemble learning can harness their unique learning capabilities more effectively compared to individual layer-wise optimization.
title PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.02122