PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Tzu-Han, Wang, How-Shing, Weng, Hao-Yung, Peng, Kuang-Chen, Chen, Zih-Ching, Lee, Hung-yi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913225469067264
author Lin, Tzu-Han
Wang, How-Shing
Weng, Hao-Yung
Peng, Kuang-Chen
Chen, Zih-Ching
Lee, Hung-yi
author_facet Lin, Tzu-Han
Wang, How-Shing
Weng, Hao-Yung
Peng, Kuang-Chen
Chen, Zih-Ching
Lee, Hung-yi
contents Parameter-Efficient Fine-Tuning (PEFT) is increasingly recognized as an effective method in speech processing. However, the optimal approach and the placement of PEFT methods remain inconclusive. Our study conducts extensive experiments to compare different PEFT methods and their layer-wise placement adapting Differentiable Architecture Search (DARTS). We also explore the use of ensemble learning to leverage diverse PEFT strategies. The results reveal that DARTS does not outperform the baseline approach, which involves inserting the same PEFT method into all layers of a Self-Supervised Learning (SSL) model. In contrast, an ensemble learning approach, particularly one employing majority voting, demonstrates superior performance. Our statistical evidence indicates that different PEFT methods learn in varied ways. This variation might explain why the synergistic integration of various PEFT methods through ensemble learning can harness their unique learning capabilities more effectively compared to individual layer-wise optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02122
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
Lin, Tzu-Han
Wang, How-Shing
Weng, Hao-Yung
Peng, Kuang-Chen
Chen, Zih-Ching
Lee, Hung-yi
Computation and Language
Sound
Audio and Speech Processing
Parameter-Efficient Fine-Tuning (PEFT) is increasingly recognized as an effective method in speech processing. However, the optimal approach and the placement of PEFT methods remain inconclusive. Our study conducts extensive experiments to compare different PEFT methods and their layer-wise placement adapting Differentiable Architecture Search (DARTS). We also explore the use of ensemble learning to leverage diverse PEFT strategies. The results reveal that DARTS does not outperform the baseline approach, which involves inserting the same PEFT method into all layers of a Self-Supervised Learning (SSL) model. In contrast, an ensemble learning approach, particularly one employing majority voting, demonstrates superior performance. Our statistical evidence indicates that different PEFT methods learn in varied ways. This variation might explain why the synergistic integration of various PEFT methods through ensemble learning can harness their unique learning capabilities more effectively compared to individual layer-wise optimization.
title PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.02122