RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Mochu, Shen, Zhelun, Li, Xuesong, Ren, Jiahui, Zhang, Jing, Zhao, Chen, Liu, Shanshan, Feng, Haocheng, Wang, Jingdong, Dai, Yuchao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914361562365952
author Xiang, Mochu
Shen, Zhelun
Li, Xuesong
Ren, Jiahui
Zhang, Jing
Zhao, Chen
Liu, Shanshan
Feng, Haocheng
Wang, Jingdong
Dai, Yuchao
author_facet Xiang, Mochu
Shen, Zhelun
Li, Xuesong
Ren, Jiahui
Zhang, Jing
Zhao, Chen
Liu, Shanshan
Feng, Haocheng
Wang, Jingdong
Dai, Yuchao
contents Human perceive the 3D world through 2D observations from limited viewpoints. While recent feed-forward generalizable 3D reconstruction models excel at recovering 3D structures from sparse images, their representations are often confined to observed regions, leaving unseen geometry un-modeled. This raises a key, fundamental challenge: Can we infer a complete 3D structure from partial 2D observations? We present RnG (Reconstruction and Generation), a novel feed-forward Transformer that unifies these two tasks by predicting an implicit, complete 3D representation. At the core of RnG, we propose a reconstruction-guided causal attention mechanism that separates reconstruction and generation at the attention level, and treats the KV-cache as an implicit 3D representation. Then, arbitrary poses can efficiently query this cache to render high-fidelity, novel-view RGBD outputs. As a result, RnG not only accurately reconstructs visible geometry but also generates plausible, coherent unseen geometry and appearance. Our method achieves state-of-the-art performance in both generalizable 3D reconstruction and novel view generation, while operating efficiently enough for real-time interactive applications. Project page: https://npucvr.github.io/RnG
format Preprint
id arxiv_https___arxiv_org_abs_2603_01194
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
Xiang, Mochu
Shen, Zhelun
Li, Xuesong
Ren, Jiahui
Zhang, Jing
Zhao, Chen
Liu, Shanshan
Feng, Haocheng
Wang, Jingdong
Dai, Yuchao
Computer Vision and Pattern Recognition
Human perceive the 3D world through 2D observations from limited viewpoints. While recent feed-forward generalizable 3D reconstruction models excel at recovering 3D structures from sparse images, their representations are often confined to observed regions, leaving unseen geometry un-modeled. This raises a key, fundamental challenge: Can we infer a complete 3D structure from partial 2D observations? We present RnG (Reconstruction and Generation), a novel feed-forward Transformer that unifies these two tasks by predicting an implicit, complete 3D representation. At the core of RnG, we propose a reconstruction-guided causal attention mechanism that separates reconstruction and generation at the attention level, and treats the KV-cache as an implicit 3D representation. Then, arbitrary poses can efficiently query this cache to render high-fidelity, novel-view RGBD outputs. As a result, RnG not only accurately reconstructs visible geometry but also generates plausible, coherent unseen geometry and appearance. Our method achieves state-of-the-art performance in both generalizable 3D reconstruction and novel view generation, while operating efficiently enough for real-time interactive applications. Project page: https://npucvr.github.io/RnG
title RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.01194