MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jiapeng, Tian, Changxin, Chen, Kunlong, Liu, Ziqi, Mao, Jiaxin, Zhao, Wayne Xin, Zhang, Zhiqiang, Zhou, Jun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914278152339456
author Wang, Jiapeng
Tian, Changxin
Chen, Kunlong
Liu, Ziqi
Mao, Jiaxin
Zhao, Wayne Xin
Zhang, Zhiqiang
Zhou, Jun
author_facet Wang, Jiapeng
Tian, Changxin
Chen, Kunlong
Liu, Ziqi
Mao, Jiaxin
Zhao, Wayne Xin
Zhang, Zhiqiang
Zhou, Jun
contents Optimizing data mixtures is essential for unlocking the full potential of large language models (LLMs), yet identifying the optimal composition remains computationally prohibitive due to reliance on heuristic trials or expensive proxy training. To address this, we introduce \textbf{MergeMix}, a novel approach that efficiently determines optimal data mixing ratios by repurposing model merging weights as a high-fidelity, low-cost performance proxy. By training domain-specific experts on minimal tokens and optimizing their merging weights against downstream benchmarks, MergeMix effectively optimizes the performance of data mixtures without incurring the cost of full-scale training. Extensive experiments on models with 8B and 16B parameters validate that MergeMix achieves performance comparable to or surpassing exhaustive manual tuning while drastically reducing search costs. Furthermore, MergeMix exhibits high rank consistency (Spearman $ρ> 0.9$) and strong cross-scale transferability, offering a scalable, automated solution for data mixture optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17858
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
Wang, Jiapeng
Tian, Changxin
Chen, Kunlong
Liu, Ziqi
Mao, Jiaxin
Zhao, Wayne Xin
Zhang, Zhiqiang
Zhou, Jun
Machine Learning
Artificial Intelligence
Optimizing data mixtures is essential for unlocking the full potential of large language models (LLMs), yet identifying the optimal composition remains computationally prohibitive due to reliance on heuristic trials or expensive proxy training. To address this, we introduce \textbf{MergeMix}, a novel approach that efficiently determines optimal data mixing ratios by repurposing model merging weights as a high-fidelity, low-cost performance proxy. By training domain-specific experts on minimal tokens and optimizing their merging weights against downstream benchmarks, MergeMix effectively optimizes the performance of data mixtures without incurring the cost of full-scale training. Extensive experiments on models with 8B and 16B parameters validate that MergeMix achieves performance comparable to or surpassing exhaustive manual tuning while drastically reducing search costs. Furthermore, MergeMix exhibits high rank consistency (Spearman $ρ> 0.9$) and strong cross-scale transferability, offering a scalable, automated solution for data mixture optimization.
title MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.17858