ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Min, Yunhong, Choi, Daehyeon, Yeo, Kyeongmin, Lee, Jihyun, Sung, Minhyuk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915578249216000
author Min, Yunhong
Choi, Daehyeon
Yeo, Kyeongmin
Lee, Jihyun
Sung, Minhyuk
author_facet Min, Yunhong
Choi, Daehyeon
Yeo, Kyeongmin
Lee, Jihyun
Sung, Minhyuk
contents We introduce ORIGEN, the first zero-shot method for 3D orientation grounding in text-to-image generation across multiple objects and diverse categories. While previous work on spatial grounding in image generation has mainly focused on 2D positioning, it lacks control over 3D orientation. To address this, we propose a reward-guided sampling approach using a pretrained discriminative model for 3D orientation estimation and a one-step text-to-image generative flow model. While gradient-ascent-based optimization is a natural choice for reward-based guidance, it struggles to maintain image realism. Instead, we adopt a sampling-based approach using Langevin dynamics, which extends gradient ascent by simply injecting random noise--requiring just a single additional line of code. Additionally, we introduce adaptive time rescaling based on the reward function to accelerate convergence. Our experiments show that ORIGEN outperforms both training-based and test-time guidance methods across quantitative metrics and user studies.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22194
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
Min, Yunhong
Choi, Daehyeon
Yeo, Kyeongmin
Lee, Jihyun
Sung, Minhyuk
Computer Vision and Pattern Recognition
Machine Learning
We introduce ORIGEN, the first zero-shot method for 3D orientation grounding in text-to-image generation across multiple objects and diverse categories. While previous work on spatial grounding in image generation has mainly focused on 2D positioning, it lacks control over 3D orientation. To address this, we propose a reward-guided sampling approach using a pretrained discriminative model for 3D orientation estimation and a one-step text-to-image generative flow model. While gradient-ascent-based optimization is a natural choice for reward-based guidance, it struggles to maintain image realism. Instead, we adopt a sampling-based approach using Langevin dynamics, which extends gradient ascent by simply injecting random noise--requiring just a single additional line of code. Additionally, we introduce adaptive time rescaling based on the reward function to accelerate convergence. Our experiments show that ORIGEN outperforms both training-based and test-time guidance methods across quantitative metrics and user studies.
title ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.22194