Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918429302194176 |
|---|---|
| author | Yin, Shaofeng Ge, Jiaxin Wang, Zora Zhiruo Wang, Chenyang Li, Xiuyu Black, Michael J. Darrell, Trevor Kanazawa, Angjoo Feng, Haiwen |
| author_facet | Yin, Shaofeng Ge, Jiaxin Wang, Zora Zhiruo Wang, Chenyang Li, Xiuyu Black, Michael J. Darrell, Trevor Kanazawa, Angjoo Feng, Haiwen |
| contents | Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained spatial grounding in one-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross-verify each other. VIGA operates through a tightly coupled code-render-inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high-level semantic skills and an evolving multimodal memory, VIGA sustains evidence-based modifications over long horizons. This training-free, task-agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi-step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual-to-code benchmark. Empirically, VIGA substantially improves accuracy compared with one-shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_11109 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning Yin, Shaofeng Ge, Jiaxin Wang, Zora Zhiruo Wang, Chenyang Li, Xiuyu Black, Michael J. Darrell, Trevor Kanazawa, Angjoo Feng, Haiwen Computer Vision and Pattern Recognition Artificial Intelligence Graphics Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained spatial grounding in one-shot settings. To address this, we introduce VIGA (Vision-as-Inverse-Graphics Agent), an interleaved multimodal reasoning framework where symbolic logic and visual perception actively cross-verify each other. VIGA operates through a tightly coupled code-render-inspect loop: synthesizing symbolic programs, projecting them into visual states, and inspecting discrepancies to guide iterative edits. Equipped with high-level semantic skills and an evolving multimodal memory, VIGA sustains evidence-based modifications over long horizons. This training-free, task-agnostic framework seamlessly supports 2D document generation, 3D reconstruction, multi-step 3D editing, and 4D physical interaction. Finally, we introduce BlenderBench, a challenging visual-to-code benchmark. Empirically, VIGA substantially improves accuracy compared with one-shot baselines in BlenderGym (35.32%), SlideBench (117.17%) and our proposed BlenderBench (124.70%). |
| title | Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Graphics |
| url | https://arxiv.org/abs/2601.11109 |