Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Guangzhao, Luo, Rundong, Ma, Wei-Chiu, Averbuch-Elor, Hadar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917556576583680
author He, Guangzhao
Luo, Rundong
Ma, Wei-Chiu
Averbuch-Elor, Hadar
author_facet He, Guangzhao
Luo, Rundong
Ma, Wei-Chiu
Averbuch-Elor, Hadar
contents Inverse graphics is a longstanding and highly underconstrained problem that seeks to reconstruct images as editable 3D scenes which can be rendered, relit, and manipulated. In this work, we investigate whether pretrained vision-language models (VLMs) can perform executable inverse graphics directly from a single image by reconstructing a scene as an editable Blender program, without relying on specialized 2D or 3D foundation models, differentiable rendering, or multi-view supervision. We introduce Staged Executable Inverse Graphics (SEIG), an agentic framework that reconstructs a 3D scene from a single image by progressively refining scene factors including geometry, materials, composition, and lighting directly in executable Blender code space. We evaluate our framework across diverse scenes using a range of reconstruction metrics spanning pixel-level, perceptual, and semantic fidelity. Our experiments show that staged reconstruction substantially improves reconstruction fidelity, highlighting the importance of task decomposition for executable inverse graphics with general-purpose VLMs. Finally, we showcase various downstream applications enabled by the reconstructed editable Blender scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2606_02580
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
He, Guangzhao
Luo, Rundong
Ma, Wei-Chiu
Averbuch-Elor, Hadar
Computer Vision and Pattern Recognition
Inverse graphics is a longstanding and highly underconstrained problem that seeks to reconstruct images as editable 3D scenes which can be rendered, relit, and manipulated. In this work, we investigate whether pretrained vision-language models (VLMs) can perform executable inverse graphics directly from a single image by reconstructing a scene as an editable Blender program, without relying on specialized 2D or 3D foundation models, differentiable rendering, or multi-view supervision. We introduce Staged Executable Inverse Graphics (SEIG), an agentic framework that reconstructs a 3D scene from a single image by progressively refining scene factors including geometry, materials, composition, and lighting directly in executable Blender code space. We evaluate our framework across diverse scenes using a range of reconstruction metrics spanning pixel-level, perceptual, and semantic fidelity. Our experiments show that staged reconstruction substantially improves reconstruction fidelity, highlighting the importance of task decomposition for executable inverse graphics with general-purpose VLMs. Finally, we showcase various downstream applications enabled by the reconstructed editable Blender scenes.
title Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.02580