MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Runjia, Sun, Guangyan, Wang, Qifan, Geng, Tong, Dianat, Sohail, Han, Xiaotian, Rao, Raghuveer, Zhang, Xueling, Han, Cheng, Huang, Lifu, Liu, Dongfang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914035158482944
author Zeng, Runjia
Sun, Guangyan
Wang, Qifan
Geng, Tong
Dianat, Sohail
Han, Xiaotian
Rao, Raghuveer
Zhang, Xueling
Han, Cheng
Huang, Lifu
Liu, Dongfang
author_facet Zeng, Runjia
Sun, Guangyan
Wang, Qifan
Geng, Tong
Dianat, Sohail
Han, Xiaotian
Rao, Raghuveer
Zhang, Xueling
Han, Cheng
Huang, Lifu
Liu, Dongfang
contents Considering deep neural networks as manifold mappers, the pretrain-then-fine-tune paradigm can be interpreted as a two-stage process: pretrain establishes a broad knowledge base, and fine-tune adjusts the model parameters to activate specific neural pathways to align with the target manifold. Although prior fine-tuning approaches demonstrate success, their rigid parameter space limits their ability to dynamically activate appropriate neural pathways, rendering them ill-equipped to adapt flexibly to the diverse and evolving data distributions. In light of this view, we propose a novel approach, Mixture of Expert Prompt Tuning (MEPT), as an effective and efficient manifold-mapping framework. MEPT leverages the Mixture of Experts architecture by integrating multiple prompt experts to adaptively learn diverse and non-stationary data distributions. Empirical evaluations demonstrate that MEPT outperforms several state-of-the-art parameter efficient baselines on SuperGLUE, achieving notable improvements in mean accuracy (e.g., 1.94%) while significantly reducing activated prompts by 79.25%. The effectiveness of MEPT is further supported by theoretical insights from manifold learning and validated through neural activation pathway visualization results. Our code is avaliable at https://runjia.tech/emnlp_mept/.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
Zeng, Runjia
Sun, Guangyan
Wang, Qifan
Geng, Tong
Dianat, Sohail
Han, Xiaotian
Rao, Raghuveer
Zhang, Xueling
Han, Cheng
Huang, Lifu
Liu, Dongfang
Machine Learning
Artificial Intelligence
Computation and Language
Considering deep neural networks as manifold mappers, the pretrain-then-fine-tune paradigm can be interpreted as a two-stage process: pretrain establishes a broad knowledge base, and fine-tune adjusts the model parameters to activate specific neural pathways to align with the target manifold. Although prior fine-tuning approaches demonstrate success, their rigid parameter space limits their ability to dynamically activate appropriate neural pathways, rendering them ill-equipped to adapt flexibly to the diverse and evolving data distributions. In light of this view, we propose a novel approach, Mixture of Expert Prompt Tuning (MEPT), as an effective and efficient manifold-mapping framework. MEPT leverages the Mixture of Experts architecture by integrating multiple prompt experts to adaptively learn diverse and non-stationary data distributions. Empirical evaluations demonstrate that MEPT outperforms several state-of-the-art parameter efficient baselines on SuperGLUE, achieving notable improvements in mean accuracy (e.g., 1.94%) while significantly reducing activated prompts by 79.25%. The effectiveness of MEPT is further supported by theoretical insights from manifold learning and validated through neural activation pathway visualization results. Our code is avaliable at https://runjia.tech/emnlp_mept/.
title MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.00996