ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cheng, Freeman, Ye, Botao, Li, Xueting, You, Junqi, Zhan, Fangneng, Yang, Ming-Hsuan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917330324291584
author Cheng, Freeman
Ye, Botao
Li, Xueting
You, Junqi
Zhan, Fangneng
Yang, Ming-Hsuan
author_facet Cheng, Freeman
Ye, Botao
Li, Xueting
You, Junqi
Zhan, Fangneng
Yang, Ming-Hsuan
contents Online novel view synthesis remains challenging, requiring robust scene reconstruction from sequential, often unposed, observations. We present ReCoSplat, an autoregressive feed-forward Gaussian Splatting model supporting posed or unposed inputs, with or without camera intrinsics. While assembling local Gaussians using camera poses scales better than canonical-space prediction, it creates a dilemma during training: using ground-truth poses ensures stability but causes a distribution mismatch when predicted poses are used at inference. To address this, we introduce a Render-and-Compare (ReCo) module. ReCo renders the current reconstruction from the predicted viewpoint and compares it with the incoming observation, providing a stable conditioning signal that compensates for pose errors. To support long sequences, we propose a hybrid KV cache compression strategy combining early-layer truncation with chunk-level selective retention, reducing the KV cache size by over 90% for 100+ frames. ReCoSplat achieves state-of-the-art performance across different input settings on both in- and out-of-distribution benchmarks. Code and pretrained models will be released. Our project page is at https://freemancheng.com/ReCoSplat .
format Preprint
id arxiv_https___arxiv_org_abs_2603_09968
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare
Cheng, Freeman
Ye, Botao
Li, Xueting
You, Junqi
Zhan, Fangneng
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
Online novel view synthesis remains challenging, requiring robust scene reconstruction from sequential, often unposed, observations. We present ReCoSplat, an autoregressive feed-forward Gaussian Splatting model supporting posed or unposed inputs, with or without camera intrinsics. While assembling local Gaussians using camera poses scales better than canonical-space prediction, it creates a dilemma during training: using ground-truth poses ensures stability but causes a distribution mismatch when predicted poses are used at inference. To address this, we introduce a Render-and-Compare (ReCo) module. ReCo renders the current reconstruction from the predicted viewpoint and compares it with the incoming observation, providing a stable conditioning signal that compensates for pose errors. To support long sequences, we propose a hybrid KV cache compression strategy combining early-layer truncation with chunk-level selective retention, reducing the KV cache size by over 90% for 100+ frames. ReCoSplat achieves state-of-the-art performance across different input settings on both in- and out-of-distribution benchmarks. Code and pretrained models will be released. Our project page is at https://freemancheng.com/ReCoSplat .
title ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.09968