Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Leng, Jiaqi, Tu, Shuyuan, Cao, Haidong, Xie, Sicheng, Dong, Daoguo, Wu, Zuxuan, Jiang, Yu-Gang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915827550257152
author Leng, Jiaqi
Tu, Shuyuan
Cao, Haidong
Xie, Sicheng
Dong, Daoguo
Wu, Zuxuan
Jiang, Yu-Gang
author_facet Leng, Jiaqi
Tu, Shuyuan
Cao, Haidong
Xie, Sicheng
Dong, Daoguo
Wu, Zuxuan
Jiang, Yu-Gang
contents Human preference alignment presents a critical yet underexplored challenge for diffusion models in text-to-3D generation. Existing solutions typically require task-specific fine-tuning, posing significant hurdles in data-scarce 3D domains. To address this, we propose Preference Score Distillation (PSD), an optimization-based framework that leverages pretrained 2D reward models for human-aligned text-to-3D synthesis without 3D training data. Our key insight stems from the incompatibility of pixel-level gradients: due to the absence of noisy samples during reward model training, direct application of 2D reward gradients disturbs the denoising process. Noticing that similar issue occurs in the naive classifier guidance in conditioned diffusion models, we fundamentally rethink preference alignment as a classifier-free guidance (CFG)-style mechanism through our implicit reward model. Furthermore, recognizing that frozen pretrained diffusion models constrain performance, we introduce an adaptive strategy to co-optimize preference scores and negative text embeddings. By incorporating CFG during optimization, online refinement of negative text embeddings dynamically enhances alignment. To our knowledge, we are the first to bridge human preference alignment with CFG theory under score distillation framework. Experiments demonstrate the superiority of PSD in aesthetic metrics, seamless integration with diverse pipelines, and strong extensibility.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01594
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
Leng, Jiaqi
Tu, Shuyuan
Cao, Haidong
Xie, Sicheng
Dong, Daoguo
Wu, Zuxuan
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Human preference alignment presents a critical yet underexplored challenge for diffusion models in text-to-3D generation. Existing solutions typically require task-specific fine-tuning, posing significant hurdles in data-scarce 3D domains. To address this, we propose Preference Score Distillation (PSD), an optimization-based framework that leverages pretrained 2D reward models for human-aligned text-to-3D synthesis without 3D training data. Our key insight stems from the incompatibility of pixel-level gradients: due to the absence of noisy samples during reward model training, direct application of 2D reward gradients disturbs the denoising process. Noticing that similar issue occurs in the naive classifier guidance in conditioned diffusion models, we fundamentally rethink preference alignment as a classifier-free guidance (CFG)-style mechanism through our implicit reward model. Furthermore, recognizing that frozen pretrained diffusion models constrain performance, we introduce an adaptive strategy to co-optimize preference scores and negative text embeddings. By incorporating CFG during optimization, online refinement of negative text embeddings dynamically enhances alignment. To our knowledge, we are the first to bridge human preference alignment with CFG theory under score distillation framework. Experiments demonstrate the superiority of PSD in aesthetic metrics, seamless integration with diverse pipelines, and strong extensibility.
title Preference Score Distillation: Leveraging 2D Rewards to Align Text-to-3D Generation with Human Preference
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.01594