Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Runlin, Chen, Yuzhao, Jain, Rahul, Hu, Yichen, Shi, Jingyu, Ramani, Karthik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909731015098368
author Duan, Runlin
Chen, Yuzhao
Jain, Rahul
Hu, Yichen
Shi, Jingyu
Ramani, Karthik
author_facet Duan, Runlin
Chen, Yuzhao
Jain, Rahul
Hu, Yichen
Shi, Jingyu
Ramani, Karthik
contents Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap, we propose Canvas3D, an interactive system leveraging a 3D engine to enable precise spatial manipulation for image generation. Upon user prompt, Canvas3D automatically converts textual descriptions into interactive objects within a 3D engine-driven virtual canvas, empowering direct and precise spatial configuration. These user-defined arrangements generate explicit spatial constraints that guide generative models in accurately reflecting user intentions in the resulting images. We conducted a closed-end comparative study between Canvas3D and a baseline system. And an open-ended study to evaluate our system "in the wild". The result indicates that Canvas3D outperforms the baseline on spatial control, interactivity, and overall user experience.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
Duan, Runlin
Chen, Yuzhao
Jain, Rahul
Hu, Yichen
Shi, Jingyu
Ramani, Karthik
Human-Computer Interaction
Generative AI (GenAI) has significantly advanced the ease and flexibility of image creation. However, it remains a challenge to precisely control spatial compositions, including object arrangement and scene conditions. To bridge this gap, we propose Canvas3D, an interactive system leveraging a 3D engine to enable precise spatial manipulation for image generation. Upon user prompt, Canvas3D automatically converts textual descriptions into interactive objects within a 3D engine-driven virtual canvas, empowering direct and precise spatial configuration. These user-defined arrangements generate explicit spatial constraints that guide generative models in accurately reflecting user intentions in the resulting images. We conducted a closed-end comparative study between Canvas3D and a baseline system. And an open-ended study to evaluate our system "in the wild". The result indicates that Canvas3D outperforms the baseline on spatial control, interactivity, and overall user experience.
title Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
topic Human-Computer Interaction
url https://arxiv.org/abs/2508.07135