Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Xingyu, Li, Xiner, Uehara, Masatoshi, Kim, Sunwoo, Zhao, Yulai, Scalia, Gabriele, Hajiramezanali, Ehsan, Biancalani, Tommaso, Zhi, Degui, Ji, Shuiwang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914360064999424
author Su, Xingyu
Li, Xiner
Uehara, Masatoshi
Kim, Sunwoo
Zhao, Yulai
Scalia, Gabriele
Hajiramezanali, Ehsan
Biancalani, Tommaso
Zhi, Degui
Ji, Shuiwang
author_facet Su, Xingyu
Li, Xiner
Uehara, Masatoshi
Kim, Sunwoo
Zhao, Yulai
Scalia, Gabriele
Hajiramezanali, Ehsan
Biancalani, Tommaso
Zhi, Degui
Ji, Shuiwang
contents We address the problem of fine-tuning diffusion models for reward-guided generation in biomolecular design. While diffusion models have proven highly effective in modeling complex, high-dimensional data distributions, real-world applications often demand more than high-fidelity generation, requiring optimization with respect to potentially non-differentiable reward functions such as physics-based simulation or rewards based on scientific knowledge. Although RL methods have been explored to fine-tune diffusion models for such objectives, they often suffer from instability, low sample efficiency, and mode collapse due to their on-policy nature. In this work, we propose an iterative distillation-based fine-tuning framework that enables diffusion models to optimize for arbitrary reward functions. Our method casts the problem as policy distillation: it collects off-policy data during the roll-in phase, simulates reward-based soft-optimal policies during roll-out, and updates the model by minimizing the KL divergence between the simulated soft-optimal policy and the current model policy. Our off-policy formulation, combined with KL divergence minimization, enhances training stability and sample efficiency compared to existing RL-based methods. Empirical results demonstrate the effectiveness and superior reward optimization of our approach across diverse tasks in protein, small molecule, and regulatory DNA design. The source code is released at (https://divelab.github.io/VIDD/).
format Preprint
id arxiv_https___arxiv_org_abs_2507_00445
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
Su, Xingyu
Li, Xiner
Uehara, Masatoshi
Kim, Sunwoo
Zhao, Yulai
Scalia, Gabriele
Hajiramezanali, Ehsan
Biancalani, Tommaso
Zhi, Degui
Ji, Shuiwang
Machine Learning
Artificial Intelligence
Quantitative Methods
We address the problem of fine-tuning diffusion models for reward-guided generation in biomolecular design. While diffusion models have proven highly effective in modeling complex, high-dimensional data distributions, real-world applications often demand more than high-fidelity generation, requiring optimization with respect to potentially non-differentiable reward functions such as physics-based simulation or rewards based on scientific knowledge. Although RL methods have been explored to fine-tune diffusion models for such objectives, they often suffer from instability, low sample efficiency, and mode collapse due to their on-policy nature. In this work, we propose an iterative distillation-based fine-tuning framework that enables diffusion models to optimize for arbitrary reward functions. Our method casts the problem as policy distillation: it collects off-policy data during the roll-in phase, simulates reward-based soft-optimal policies during roll-out, and updates the model by minimizing the KL divergence between the simulated soft-optimal policy and the current model policy. Our off-policy formulation, combined with KL divergence minimization, enhances training stability and sample efficiency compared to existing RL-based methods. Empirical results demonstrate the effectiveness and superior reward optimization of our approach across diverse tasks in protein, small molecule, and regulatory DNA design. The source code is released at (https://divelab.github.io/VIDD/).
title Iterative Distillation for Reward-Guided Fine-Tuning of Diffusion Models in Biomolecular Design
topic Machine Learning
Artificial Intelligence
Quantitative Methods
url https://arxiv.org/abs/2507.00445