Timber: Training-free Instruct Model Refining with Base via Effective Rank

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Taiqiang, Yang, Runming, Liu, Tao, Wang, Jiahao, Xu, Zenan, Wong, Ngai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912612344659968
author Wu, Taiqiang
Yang, Runming
Liu, Tao
Wang, Jiahao
Xu, Zenan
Wong, Ngai
author_facet Wu, Taiqiang
Yang, Runming
Liu, Tao
Wang, Jiahao
Xu, Zenan
Wong, Ngai
contents Post-training, which elicits a pretrained Base model into the corresponding Instruct model, is widely considered to be superficial. In this work, we first reinforce this hypothesis by providing novel quantitative evidence from the weight level that the effective rank (eRank) remains negligibly changed. However, this superficiality also suffers a critical trade-off, improving the exploitation capabilities at the cost of limiting its exploration. To tackle this issue, we propose Timber, a simple yet effective training-free method that enhances the exploration capability of the Instruct model while preserving its exploitation. The key insight is to partially revert Instruct towards the paired Base model by subtle yet targeted refinement of the weight deltas. Extensive experiments on Llama and Qwen series demonstrate that Timber consistently improves vanilla Instruct models, particularly on Pass@k performance. Our findings offer new insights into the post-training stage at the weight level and practical strategies to refine the Instruct model without training.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23595
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Timber: Training-free Instruct Model Refining with Base via Effective Rank
Wu, Taiqiang
Yang, Runming
Liu, Tao
Wang, Jiahao
Xu, Zenan
Wong, Ngai
Computation and Language
Artificial Intelligence
Post-training, which elicits a pretrained Base model into the corresponding Instruct model, is widely considered to be superficial. In this work, we first reinforce this hypothesis by providing novel quantitative evidence from the weight level that the effective rank (eRank) remains negligibly changed. However, this superficiality also suffers a critical trade-off, improving the exploitation capabilities at the cost of limiting its exploration. To tackle this issue, we propose Timber, a simple yet effective training-free method that enhances the exploration capability of the Instruct model while preserving its exploitation. The key insight is to partially revert Instruct towards the paired Base model by subtle yet targeted refinement of the weight deltas. Extensive experiments on Llama and Qwen series demonstrate that Timber consistently improves vanilla Instruct models, particularly on Pass@k performance. Our findings offer new insights into the post-training stage at the weight level and practical strategies to refine the Instruct model without training.
title Timber: Training-free Instruct Model Refining with Base via Effective Rank
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.23595