CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Feng, Jinyuan, Wei, Chaopeng, Qiu, Tenghai, Hu, Tianyi, Pu, Zhiqiang
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916922594951168
author Feng, Jinyuan
Wei, Chaopeng
Qiu, Tenghai
Hu, Tianyi
Pu, Zhiqiang
author_facet Feng, Jinyuan
Wei, Chaopeng
Qiu, Tenghai
Hu, Tianyi
Pu, Zhiqiang
contents In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has been widely adopted as a promising approach to trade-off between model capacity and computation overhead. However, current MoE variants fall short on heterogeneous datasets, ignoring the fact that experts may learn similar knowledge, resulting in the underutilization of MoE's capacity. In this paper, we propose Contrastive Representation for MoE (CoMoE), a novel method to promote modularization and specialization in MoE, where the experts are trained along with a contrastive objective by sampling from activated and inactivated experts in top-k routing. We demonstrate that such a contrastive objective recovers the mutual-information gap between inputs and the two types of experts. Experiments on several benchmarks and in multi-task settings demonstrate that CoMoE can consistently enhance MoE's capacity and promote modularization among the experts.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
Feng, Jinyuan
Wei, Chaopeng
Qiu, Tenghai
Hu, Tianyi
Pu, Zhiqiang
Machine Learning
Computation and Language
In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has been widely adopted as a promising approach to trade-off between model capacity and computation overhead. However, current MoE variants fall short on heterogeneous datasets, ignoring the fact that experts may learn similar knowledge, resulting in the underutilization of MoE's capacity. In this paper, we propose Contrastive Representation for MoE (CoMoE), a novel method to promote modularization and specialization in MoE, where the experts are trained along with a contrastive objective by sampling from activated and inactivated experts in top-k routing. We demonstrate that such a contrastive objective recovers the mutual-information gap between inputs and the two types of experts. Experiments on several benchmarks and in multi-task settings demonstrate that CoMoE can consistently enhance MoE's capacity and promote modularization among the experts.
title CoMoE: Contrastive Representation for Mixture-of-Experts in Parameter-Efficient Fine-tuning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.17553