Déjà View: Looping Transformers for Multi-View 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burzio, Alessandro, Fischer, Tobias, Elflein, Sven, Zhou, Qunjie, de Lutio, Riccardo, Ren, Jiawei, Huang, Jiahui, Huang, Shengyu, Pollefeys, Marc, Leal-Taixé, Laura, Gojcic, Zan, Turki, Haithem
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913173493252096
author Burzio, Alessandro
Fischer, Tobias
Elflein, Sven
Zhou, Qunjie
de Lutio, Riccardo
Ren, Jiawei
Huang, Jiahui
Huang, Shengyu
Pollefeys, Marc
Leal-Taixé, Laura
Gojcic, Zan
Turki, Haithem
author_facet Burzio, Alessandro
Fischer, Tobias
Elflein, Sven
Zhou, Qunjie
de Lutio, Riccardo
Ren, Jiawei
Huang, Jiahui
Huang, Shengyu
Pollefeys, Marc
Leal-Taixé, Laura
Gojcic, Zan
Turki, Haithem
contents Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emerging evidence suggests that contiguous transformer layers often behave like repeated applications of similar operations, and multi-view reconstruction transformers refine their predictions progressively across decoder depth. We posit that model depth partially buys iteration, paid for inefficiently in unique parameters, and instead make that iteration explicit in architecture. Our model, DéjàView, applies a single looped transformer block recurrently to per-view features for K refinement steps. Trained once, it exposes K as an inference-time compute knob, matching or outperforming substantially larger feed-forward baselines across five reconstruction benchmarks spanning indoor, outdoor, object-centric, and driving scenes, while using a fraction of their parameters and comparable or lower compute. Importantly, the same looped block formulation outperforms an otherwise identical variant with independent per-step parameters under matched training data and compute, suggesting that explicit iteration is not merely a compute-efficient substitute for capacity but a stronger inductive bias for multi-view 3D reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30215
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Déjà View: Looping Transformers for Multi-View 3D Reconstruction
Burzio, Alessandro
Fischer, Tobias
Elflein, Sven
Zhou, Qunjie
de Lutio, Riccardo
Ren, Jiawei
Huang, Jiahui
Huang, Shengyu
Pollefeys, Marc
Leal-Taixé, Laura
Gojcic, Zan
Turki, Haithem
Computer Vision and Pattern Recognition
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emerging evidence suggests that contiguous transformer layers often behave like repeated applications of similar operations, and multi-view reconstruction transformers refine their predictions progressively across decoder depth. We posit that model depth partially buys iteration, paid for inefficiently in unique parameters, and instead make that iteration explicit in architecture. Our model, DéjàView, applies a single looped transformer block recurrently to per-view features for K refinement steps. Trained once, it exposes K as an inference-time compute knob, matching or outperforming substantially larger feed-forward baselines across five reconstruction benchmarks spanning indoor, outdoor, object-centric, and driving scenes, while using a fraction of their parameters and comparable or lower compute. Importantly, the same looped block formulation outperforms an otherwise identical variant with independent per-step parameters under matched training data and compute, suggesting that explicit iteration is not merely a compute-efficient substitute for capacity but a stronger inductive bias for multi-view 3D reconstruction.
title Déjà View: Looping Transformers for Multi-View 3D Reconstruction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.30215