EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Song, Wu, Fan, Zhang, Lei, Zheng, Xiawu, Zhang, Shengchuan, Chao, Fei, Shi, Yiyu, Ji, Rongrong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913237923004416
author Guo, Song
Wu, Fan
Zhang, Lei
Zheng, Xiawu
Zhang, Shengchuan
Chao, Fei
Shi, Yiyu
Ji, Rongrong
author_facet Guo, Song
Wu, Fan
Zhang, Lei
Zheng, Xiawu
Zhang, Shengchuan
Chao, Fei
Shi, Yiyu
Ji, Rongrong
contents Existing methods for fine-tuning sparse LLMs often suffer from resource-intensive requirements and high retraining costs. Additionally, many fine-tuning methods often rely on approximations or heuristic optimization strategies, which may lead to suboptimal solutions. To address these issues, we propose an efficient and fast framework for fine-tuning sparse LLMs based on minimizing reconstruction error. Our approach involves sampling a small dataset for calibration and utilizing backpropagation to iteratively optimize block-wise reconstruction error, on a block-by-block basis, aiming for optimal solutions. Extensive experiments on various benchmarks consistently demonstrate the superiority of our method over other baselines. For instance, on the Wikitext2 dataset with LlamaV1-7B at 70% sparsity, our proposed EBFT achieves a perplexity of 16.88, surpassing the state-of-the-art DSnoT with a perplexity of 75.14. Moreover, with a structured sparsity ratio of 26\%, EBFT achieves a perplexity of 16.27, outperforming LoRA (perplexity 16.44). Furthermore, the fine-tuning process of EBFT for LlamaV1-7B only takes approximately 30 minutes, and the entire framework can be executed on a single 16GB GPU. The source code is available at https://github.com/sunggo/EBFT.
format Preprint
id arxiv_https___arxiv_org_abs_2402_12419
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
Guo, Song
Wu, Fan
Zhang, Lei
Zheng, Xiawu
Zhang, Shengchuan
Chao, Fei
Shi, Yiyu
Ji, Rongrong
Machine Learning
Artificial Intelligence
Computation and Language
Existing methods for fine-tuning sparse LLMs often suffer from resource-intensive requirements and high retraining costs. Additionally, many fine-tuning methods often rely on approximations or heuristic optimization strategies, which may lead to suboptimal solutions. To address these issues, we propose an efficient and fast framework for fine-tuning sparse LLMs based on minimizing reconstruction error. Our approach involves sampling a small dataset for calibration and utilizing backpropagation to iteratively optimize block-wise reconstruction error, on a block-by-block basis, aiming for optimal solutions. Extensive experiments on various benchmarks consistently demonstrate the superiority of our method over other baselines. For instance, on the Wikitext2 dataset with LlamaV1-7B at 70% sparsity, our proposed EBFT achieves a perplexity of 16.88, surpassing the state-of-the-art DSnoT with a perplexity of 75.14. Moreover, with a structured sparsity ratio of 26\%, EBFT achieves a perplexity of 16.27, outperforming LoRA (perplexity 16.44). Furthermore, the fine-tuning process of EBFT for LlamaV1-7B only takes approximately 30 minutes, and the entire framework can be executed on a single 16GB GPU. The source code is available at https://github.com/sunggo/EBFT.
title EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2402.12419