Large Language Models to Diffusion Finetuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cetin, Edoardo, Zhao, Tianyu, Tang, Yujin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910981597167616
author Cetin, Edoardo
Zhao, Tianyu
Tang, Yujin
author_facet Cetin, Edoardo
Zhao, Tianyu
Tang, Yujin
contents We propose a new finetuning method to provide pre-trained large language models (LMs) the ability to scale test-time compute through the diffusion framework. By increasing the number of diffusion steps, we show our finetuned models achieve monotonically increasing accuracy, directly translating to improved performance across downstream tasks. Furthermore, our finetuned models can expertly answer questions on specific topics by integrating powerful guidance techniques, and autonomously determine the compute required for a given problem by leveraging adaptive ODE solvers. Our method is universally applicable to any foundation model pre-trained with a cross-entropy loss and does not modify any of its original weights, fully preserving its strong single-step generation capabilities. We show our method is more effective and fully compatible with traditional finetuning approaches, introducing an orthogonal new direction to unify the strengths of the autoregressive and diffusion frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models to Diffusion Finetuning
Cetin, Edoardo
Zhao, Tianyu
Tang, Yujin
Computation and Language
Artificial Intelligence
Machine Learning
We propose a new finetuning method to provide pre-trained large language models (LMs) the ability to scale test-time compute through the diffusion framework. By increasing the number of diffusion steps, we show our finetuned models achieve monotonically increasing accuracy, directly translating to improved performance across downstream tasks. Furthermore, our finetuned models can expertly answer questions on specific topics by integrating powerful guidance techniques, and autonomously determine the compute required for a given problem by leveraging adaptive ODE solvers. Our method is universally applicable to any foundation model pre-trained with a cross-entropy loss and does not modify any of its original weights, fully preserving its strong single-step generation capabilities. We show our method is more effective and fully compatible with traditional finetuning approaches, introducing an orthogonal new direction to unify the strengths of the autoregressive and diffusion frameworks.
title Large Language Models to Diffusion Finetuning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.15781