CAMEx: Curvature-aware Merging of Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Dung V., Nguyen, Minh H., Nguyen, Luc Q., Teo, Rachel S. Y., Nguyen, Tan M., Tran, Linh Duy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910855330791424
author Nguyen, Dung V.
Nguyen, Minh H.
Nguyen, Luc Q.
Teo, Rachel S. Y.
Nguyen, Tan M.
Tran, Linh Duy
author_facet Nguyen, Dung V.
Nguyen, Minh H.
Nguyen, Luc Q.
Teo, Rachel S. Y.
Nguyen, Tan M.
Tran, Linh Duy
contents Existing methods for merging experts during model training and fine-tuning predominantly rely on Euclidean geometry, which assumes a flat parameter space. This assumption can limit the model's generalization ability, especially during the pre-training phase, where the parameter manifold might exhibit more complex curvature. Curvature-aware merging methods typically require additional information and computational resources to approximate the Fisher Information Matrix, adding memory overhead. In this paper, we introduce CAMEx (Curvature-Aware Merging of Experts), a novel expert merging protocol that incorporates natural gradients to account for the non-Euclidean curvature of the parameter manifold. By leveraging natural gradients, CAMEx adapts more effectively to the structure of the parameter space, improving alignment between model updates and the manifold's geometry. This approach enhances both pre-training and fine-tuning, resulting in better optimization trajectories and improved generalization without the substantial memory overhead typically associated with curvature-aware methods. Our contributions are threefold: (1) CAMEx significantly outperforms traditional Euclidean-based expert merging techniques across various natural language processing tasks, leading to enhanced performance during pre-training and fine-tuning; (2) we introduce a dynamic merging architecture that optimizes resource utilization, achieving high performance while reducing computational costs, facilitating efficient scaling of large language models; and (3) we provide both theoretical and empirical evidence to demonstrate the efficiency of our proposed method. The code is publicly available at: https://github.com/kpup1710/CAMEx.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18821
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CAMEx: Curvature-aware Merging of Experts
Nguyen, Dung V.
Nguyen, Minh H.
Nguyen, Luc Q.
Teo, Rachel S. Y.
Nguyen, Tan M.
Tran, Linh Duy
Machine Learning
Existing methods for merging experts during model training and fine-tuning predominantly rely on Euclidean geometry, which assumes a flat parameter space. This assumption can limit the model's generalization ability, especially during the pre-training phase, where the parameter manifold might exhibit more complex curvature. Curvature-aware merging methods typically require additional information and computational resources to approximate the Fisher Information Matrix, adding memory overhead. In this paper, we introduce CAMEx (Curvature-Aware Merging of Experts), a novel expert merging protocol that incorporates natural gradients to account for the non-Euclidean curvature of the parameter manifold. By leveraging natural gradients, CAMEx adapts more effectively to the structure of the parameter space, improving alignment between model updates and the manifold's geometry. This approach enhances both pre-training and fine-tuning, resulting in better optimization trajectories and improved generalization without the substantial memory overhead typically associated with curvature-aware methods. Our contributions are threefold: (1) CAMEx significantly outperforms traditional Euclidean-based expert merging techniques across various natural language processing tasks, leading to enhanced performance during pre-training and fine-tuning; (2) we introduce a dynamic merging architecture that optimizes resource utilization, achieving high performance while reducing computational costs, facilitating efficient scaling of large language models; and (3) we provide both theoretical and empirical evidence to demonstrate the efficiency of our proposed method. The code is publicly available at: https://github.com/kpup1710/CAMEx.
title CAMEx: Curvature-aware Merging of Experts
topic Machine Learning
url https://arxiv.org/abs/2502.18821