Revisiting Model Interpolation for Efficient Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Taiqiang, Yang, Runming, Liu, Tao, Wang, Jiahao, Wong, Ngai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918304455589888
author Wu, Taiqiang
Yang, Runming
Liu, Tao
Wang, Jiahao
Wong, Ngai
author_facet Wu, Taiqiang
Yang, Runming
Liu, Tao
Wang, Jiahao
Wong, Ngai
contents Model merging, typically on Instruct and Thinking models, has shown remarkable performance for efficient reasoning. In this paper, we systematically revisit the simplest merging method that interpolates two weights directly. Particularly, we observe that model interpolation follows a three-stage evolutionary paradigm with distinct behaviors on the reasoning trajectory. These dynamics provide a principled guide for navigating the performance-cost trade-off. Empirical results demonstrate that a strategically interpolated model surprisingly surpasses sophisticated model merging baselines on both efficiency and effectiveness. We further validate our findings with extensive ablation studies on model layers, modules, and decoding strategies. Ultimately, this work demystifies model interpolation and offers a practical framework for crafting models with precisely targeted reasoning capabilities. Code is available at \href{https://github.com/wutaiqiang/MI}{Github}.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10977
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Model Interpolation for Efficient Reasoning
Wu, Taiqiang
Yang, Runming
Liu, Tao
Wang, Jiahao
Wong, Ngai
Artificial Intelligence
Computation and Language
Model merging, typically on Instruct and Thinking models, has shown remarkable performance for efficient reasoning. In this paper, we systematically revisit the simplest merging method that interpolates two weights directly. Particularly, we observe that model interpolation follows a three-stage evolutionary paradigm with distinct behaviors on the reasoning trajectory. These dynamics provide a principled guide for navigating the performance-cost trade-off. Empirical results demonstrate that a strategically interpolated model surprisingly surpasses sophisticated model merging baselines on both efficiency and effectiveness. We further validate our findings with extensive ablation studies on model layers, modules, and decoding strategies. Ultimately, this work demystifies model interpolation and offers a practical framework for crafting models with precisely targeted reasoning capabilities. Code is available at \href{https://github.com/wutaiqiang/MI}{Github}.
title Revisiting Model Interpolation for Efficient Reasoning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.10977