Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914412190760960 |
|---|---|
| author | Zhou, Yujie Ling, Pengyang Bu, Jiazi Gao, Bingjie Niu, Li |
| author_facet | Zhou, Yujie Ling, Pengyang Bu, Jiazi Gao, Bingjie Niu, Li |
| contents | In practical AI workflows, complex tasks often involve chaining multiple generative models, such as using a video or 3D generation model after a 2D image generator. However, distributional mismatches between the output of upstream models and the expected input of downstream models frequently degrade overall generation quality. To address this issue, we propose Uni-Classifier (Uni-C), a simple yet effective plug-and-play module that leverages video diffusion priors to guide the denoising process of preceding models, thereby aligning their outputs with downstream requirements. Uni-C can also be applied independently to enhance the output quality of individual generative models. Extensive experiments across video and 3D generation tasks demonstrate that Uni-C consistently improves generation quality in both workflow-based and standalone settings, highlighting its versatility and strong generalization capability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_20382 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier Zhou, Yujie Ling, Pengyang Bu, Jiazi Gao, Bingjie Niu, Li Computer Vision and Pattern Recognition In practical AI workflows, complex tasks often involve chaining multiple generative models, such as using a video or 3D generation model after a 2D image generator. However, distributional mismatches between the output of upstream models and the expected input of downstream models frequently degrade overall generation quality. To address this issue, we propose Uni-Classifier (Uni-C), a simple yet effective plug-and-play module that leverages video diffusion priors to guide the denoising process of preceding models, thereby aligning their outputs with downstream requirements. Uni-C can also be applied independently to enhance the output quality of individual generative models. Extensive experiments across video and 3D generation tasks demonstrate that Uni-C consistently improves generation quality in both workflow-based and standalone settings, highlighting its versatility and strong generalization capability. |
| title | Uni-Classifier: Leveraging Video Diffusion Priors for Universal Guidance Classifier |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.20382 |