Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Shaofeng, Ge, Jiaxin, Wang, Zora Zhiruo, Wang, Chenyang, Li, Xiuyu, Black, Michael J., Darrell, Trevor, Kanazawa, Angjoo, Feng, Haiwen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918429302194176
author Yin, Shaofeng
Ge, Jiaxin
Wang, Zora Zhiruo
Wang, Chenyang
Li, Xiuyu
Black, Michael J.
Darrell, Trevor
Kanazawa, Angjoo
Feng, Haiwen
author_facet Yin, Shaofeng
Ge, Jiaxin
Wang, Zora Zhiruo
Wang, Chenyang
Li, Xiuyu
Black, Michael J.
Darrell, Trevor
Kanazawa, Angjoo
Feng, Haiwen
contents Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained spatial grounding in one-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross-verify each other. VIGA operates through a tightly coupled code-render-inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high-level semantic skills and an evolving multimodal memory, VIGA sustains evidence-based modifications over long horizons. This training-free, task-agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi-step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual-to-code benchmark. Empirically, VIGA substantially improves accuracy compared with one-shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%).
format Preprint
id arxiv_https___arxiv_org_abs_2601_11109
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Yin, Shaofeng
Ge, Jiaxin
Wang, Zora Zhiruo
Wang, Chenyang
Li, Xiuyu
Black, Michael J.
Darrell, Trevor
Kanazawa, Angjoo
Feng, Haiwen
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained spatial grounding in one-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross-verify each other. VIGA operates through a tightly coupled code-render-inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high-level semantic skills and an evolving multimodal memory, VIGA sustains evidence-based modifications over long horizons. This training-free, task-agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi-step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual-to-code benchmark. Empirically, VIGA substantially improves accuracy compared with one-shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%).
title Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2601.11109