FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Zhongyu, Dong, Menghang, Zhang, Rongyu, Zheng, Wenzhao, Zhang, Yunpeng, Yang, Huanrui, Du, Dalong, Keutzer, Kurt, Zhang, Shanghang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911998989565952
author Zhao, Zhongyu
Dong, Menghang
Zhang, Rongyu
Zheng, Wenzhao
Zhang, Yunpeng
Yang, Huanrui
Du, Dalong
Keutzer, Kurt
Zhang, Shanghang
author_facet Zhao, Zhongyu
Dong, Menghang
Zhang, Rongyu
Zheng, Wenzhao
Zhang, Yunpeng
Yang, Huanrui
Du, Dalong
Keutzer, Kurt
Zhang, Shanghang
contents Recent research has demonstrated that Feed-Forward Networks (FFNs) in Large Language Models (LLMs) play a pivotal role in storing diverse linguistic and factual knowledge. Conventional methods frequently face challenges due to knowledge confusion stemming from their monolithic and redundant architectures, which calls for more efficient solutions with minimal computational overhead, particularly for LLMs. In this paper, we explore the FFN computation paradigm in LLMs and introduce FactorLLM, a novel approach that decomposes well-trained dense FFNs into sparse sub-networks without requiring any further modifications, while maintaining the same level of performance. Furthermore, we embed a router from the Mixture-of-Experts (MoE), combined with our devised Prior-Approximate (PA) loss term that facilitates the dynamic activation of experts and knowledge adaptation, thereby accelerating computational processes and enhancing performance using minimal training data and fine-tuning steps. FactorLLM thus enables efficient knowledge factorization and activates select groups of experts specifically tailored to designated tasks, emulating the interactive functional segmentation of the human brain. Extensive experiments across various benchmarks demonstrate the effectiveness of our proposed FactorLLM which achieves comparable performance to the source model securing up to 85% model performance while obtaining over a 30% increase in inference speed. Code: https://github.com/zhenwuweihe/FactorLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11855
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
Zhao, Zhongyu
Dong, Menghang
Zhang, Rongyu
Zheng, Wenzhao
Zhang, Yunpeng
Yang, Huanrui
Du, Dalong
Keutzer, Kurt
Zhang, Shanghang
Computation and Language
Artificial Intelligence
Machine Learning
Recent research has demonstrated that Feed-Forward Networks (FFNs) in Large Language Models (LLMs) play a pivotal role in storing diverse linguistic and factual knowledge. Conventional methods frequently face challenges due to knowledge confusion stemming from their monolithic and redundant architectures, which calls for more efficient solutions with minimal computational overhead, particularly for LLMs. In this paper, we explore the FFN computation paradigm in LLMs and introduce FactorLLM, a novel approach that decomposes well-trained dense FFNs into sparse sub-networks without requiring any further modifications, while maintaining the same level of performance. Furthermore, we embed a router from the Mixture-of-Experts (MoE), combined with our devised Prior-Approximate (PA) loss term that facilitates the dynamic activation of experts and knowledge adaptation, thereby accelerating computational processes and enhancing performance using minimal training data and fine-tuning steps. FactorLLM thus enables efficient knowledge factorization and activates select groups of experts specifically tailored to designated tasks, emulating the interactive functional segmentation of the human brain. Extensive experiments across various benchmarks demonstrate the effectiveness of our proposed FactorLLM which achieves comparable performance to the source model securing up to 85% model performance while obtaining over a 30% increase in inference speed. Code: https://github.com/zhenwuweihe/FactorLLM.
title FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2408.11855