Complexity-aware fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Goncharov, Andrey, Vyazhev, Daniil, Sychev, Petr, Khalafyan, Edvard, Zaytsev, Alexey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917357254868992
author Goncharov, Andrey
Vyazhev, Daniil
Sychev, Petr
Khalafyan, Edvard
Zaytsev, Alexey
author_facet Goncharov, Andrey
Vyazhev, Daniil
Sychev, Petr
Khalafyan, Edvard
Zaytsev, Alexey
contents General-purpose Large Language Models (LLMs) are frequently fine-tuned through supervised fine-tuning (SFT) to enhance performance in specific domains. Better results can be achieved by distilling the chain-of-thought of a larger model at the cost of numerous expensive calls and a much greater amount of data. We propose a novel blueprint for efficient fine-tuning that uses reasoning only for complex data identified by entropy. Specifically, across three small open models ($\approx 3B$) we split the training data into complexity categories by a single token answer entropy (ROC AUC $0.73$), fine-tune large language models (LLMs) via SFT and distillation, and show that our pipeline significantly outperforms the standard SFT approach ($0.58$ vs $0.45$ average accuracy) and outperforms the distillation approach ($0.58$ vs $0.56$ average accuracy) while using $81\%$ less data.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21220
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Complexity-aware fine-tuning
Goncharov, Andrey
Vyazhev, Daniil
Sychev, Petr
Khalafyan, Edvard
Zaytsev, Alexey
Machine Learning
Computation and Language
General-purpose Large Language Models (LLMs) are frequently fine-tuned through supervised fine-tuning (SFT) to enhance performance in specific domains. Better results can be achieved by distilling the chain-of-thought of a larger model at the cost of numerous expensive calls and a much greater amount of data. We propose a novel blueprint for efficient fine-tuning that uses reasoning only for complex data identified by entropy. Specifically, across three small open models ($\approx 3B$) we split the training data into complexity categories by a single token answer entropy (ROC AUC $0.73$), fine-tune large language models (LLMs) via SFT and distillation, and show that our pipeline significantly outperforms the standard SFT approach ($0.58$ vs $0.45$ average accuracy) and outperforms the distillation approach ($0.58$ vs $0.56$ average accuracy) while using $81\%$ less data.
title Complexity-aware fine-tuning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2506.21220