Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Yujie, Ling, Pengyang, Bu, Jiazi, Gao, Bingjie, Niu, Li
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914412190760960
author Zhou, Yujie
Ling, Pengyang
Bu, Jiazi
Gao, Bingjie
Niu, Li
author_facet Zhou, Yujie
Ling, Pengyang
Bu, Jiazi
Gao, Bingjie
Niu, Li
contents In practical AI workflows, complex tasks often involve chaining multiple generative models, such as using a video or 3D generation model after a 2D image generator. However, distributional mismatches between the output of upstream models and the expected input of downstream models frequently degrade overall generation quality. To address this issue, we propose Uni-Classifier (Uni-C), a simple yet effective plug-and-play module that leverages video diffusion priors to guide the denoising process of preceding models, thereby aligning their outputs with downstream requirements. Uni-C can also be applied independently to enhance the output quality of individual generative models. Extensive experiments across video and 3D generation tasks demonstrate that Uni-C consistently improves generation quality in both workflow-based and standalone settings, highlighting its versatility and strong generalization capability.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20382
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier
Zhou, Yujie
Ling, Pengyang
Bu, Jiazi
Gao, Bingjie
Niu, Li
Computer Vision and Pattern Recognition
In practical AI workflows, complex tasks often involve chaining multiple generative models, such as using a video or 3D generation model after a 2D image generator. However, distributional mismatches between the output of upstream models and the expected input of downstream models frequently degrade overall generation quality. To address this issue, we propose Uni-Classifier (Uni-C), a simple yet effective plug-and-play module that leverages video diffusion priors to guide the denoising process of preceding models, thereby aligning their outputs with downstream requirements. Uni-C can also be applied independently to enhance the output quality of individual generative models. Extensive experiments across video and 3D generation tasks demonstrate that Uni-C consistently improves generation quality in both workflow-based and standalone settings, highlighting its versatility and strong generalization capability.
title Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.20382