D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Dengyang, Jin, Xin, Liu, Dongyang, Wang, Zanyi, Zheng, Mingzhe, Du, Ruoyi, Yang, Xiangpeng, Wu, Qilong, Li, Zhen, Gao, Peng, Yang, Harry, Hoi, Steven
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914604262621184
author Jiang, Dengyang
Jin, Xin
Liu, Dongyang
Wang, Zanyi
Zheng, Mingzhe
Du, Ruoyi
Yang, Xiangpeng
Wu, Qilong
Li, Zhen
Gao, Peng
Yang, Harry
Hoi, Steven
author_facet Jiang, Dengyang
Jin, Xin
Liu, Dongyang
Wang, Zanyi
Zheng, Mingzhe
Du, Ruoyi
Yang, Xiangpeng
Wu, Qilong
Li, Zhen
Gao, Peng
Yang, Harry
Hoi, Steven
contents The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for direct continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromise their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in-context capabilities. This enables us to formulate the training as an on-policy self-distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimizing on the model's own trajectory and under its own supervision, D-OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few-step capacity.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05204
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
Jiang, Dengyang
Jin, Xin
Liu, Dongyang
Wang, Zanyi
Zheng, Mingzhe
Du, Ruoyi
Yang, Xiangpeng
Wu, Qilong
Li, Zhen
Gao, Peng
Yang, Harry
Hoi, Steven
Computer Vision and Pattern Recognition
The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for direct continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromise their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in-context capabilities. This enables us to formulate the training as an on-policy self-distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimizing on the model's own trajectory and under its own supervision, D-OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few-step capacity.
title D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.05204