Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917480664924160 |
|---|---|
| author | Zhang, Yanbing Wang, Bo Liu, Jianhui Jiang, Nan Jiang, Jiaxiu Sun, Haoze Yang, Yijun Zheng, Shenghe Song, Lin Huang, Haoyang Duan, Nan Li, Wenbo |
| author_facet | Zhang, Yanbing Wang, Bo Liu, Jianhui Jiang, Nan Jiang, Jiaxiu Sun, Haoze Yang, Yijun Zheng, Shenghe Song, Lin Huang, Haoyang Duan, Nan Li, Wenbo |
| contents | Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static observation. We propose Thinking with Novel Views (TwNV), a paradigm that integrates generative novel-view synthesis into the reasoning loop: a Reasoner LMM identifies spatial ambiguity, instructs a Painter to synthesize an alternative viewpoint, and re-examines the scene with the additional evidence. Through systematic experiments we address three research questions. (1) Instruction format: numerical camera-pose specifications yield more reliable view control than free-form language. (2) Generation fidelity: synthesized view quality is tightly coupled with downstream spatial accuracy. (3) Inference-time visual scaling: iterative multi-turn view refinement further improves performance, echoing recent scaling trends in language reasoning. Across four spatial subtask categories and four LMM architectures (both closed- and open-source), TwNV consistently improves accuracy by +1.3 to +3.9 pp, with the largest gains on viewpoint-sensitive subtasks. These results establish novel-view generation as a practical lever for advancing spatial intelligence of LMMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_10588 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence Zhang, Yanbing Wang, Bo Liu, Jianhui Jiang, Nan Jiang, Jiaxiu Sun, Haoze Yang, Yijun Zheng, Shenghe Song, Lin Huang, Haoyang Duan, Nan Li, Wenbo Computer Vision and Pattern Recognition Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static observation. We propose Thinking with Novel Views (TwNV), a paradigm that integrates generative novel-view synthesis into the reasoning loop: a Reasoner LMM identifies spatial ambiguity, instructs a Painter to synthesize an alternative viewpoint, and re-examines the scene with the additional evidence. Through systematic experiments we address three research questions. (1) Instruction format: numerical camera-pose specifications yield more reliable view control than free-form language. (2) Generation fidelity: synthesized view quality is tightly coupled with downstream spatial accuracy. (3) Inference-time visual scaling: iterative multi-turn view refinement further improves performance, echoing recent scaling trends in language reasoning. Across four spatial subtask categories and four LMM architectures (both closed- and open-source), TwNV consistently improves accuracy by +1.3 to +3.9 pp, with the largest gains on viewpoint-sensitive subtasks. These results establish novel-view generation as a practical lever for advancing spatial intelligence of LMMs. |
| title | Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.10588 |