LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Junyi, Herrmann, Charles, Hur, Junhwa, Sun, Chen, Yang, Ming-Hsuan, Cole, Forrester, Darrell, Trevor, Sun, Deqing
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913063754530816
author Zhang, Junyi
Herrmann, Charles
Hur, Junhwa
Sun, Chen
Yang, Ming-Hsuan
Cole, Forrester
Darrell, Trevor
Sun, Deqing
author_facet Zhang, Junyi
Herrmann, Charles
Hur, Junhwa
Sun, Chen
Yang, Ming-Hsuan
Cole, Forrester
Darrell, Trevor
Sun, Deqing
contents Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or limited effective memory in recurrent designs. We present LoGeR (Long-context Geometric Reconstruction), a novel architecture that scales dense 3D reconstruction to extremely long sequences without post-optimization. LoGeR processes video streams in chunks, leveraging strong bidirectional priors for high-fidelity intra-chunk reasoning. To manage the critical challenge of coherence across chunk boundaries, we propose a learning-based hybrid memory module. This dual-component system combines a parametric Test-Time Training (TTT) memory to anchor the global coordinate frame and prevent scale drift, alongside a non-parametric Sliding Window Attention (SWA) mechanism to preserve uncompressed context for high-precision adjacent alignment. Remarkably, this memory architecture enables LoGeR to be trained on sequences of 128 frames, and generalize up to thousands of frames during inference. Evaluated across standard benchmarks and a newly repurposed VBR dataset with sequences of up to 19k frames, LoGeR substantially outperforms prior state-of-the-art feedforward methods--reducing ATE on KITTI by over 74%--and achieves robust, globally consistent reconstruction over unprecedented horizons.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03269
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Zhang, Junyi
Herrmann, Charles
Hur, Junhwa
Sun, Chen
Yang, Ming-Hsuan
Cole, Forrester
Darrell, Trevor
Sun, Deqing
Computer Vision and Pattern Recognition
Machine Learning
Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadratic attention complexity or limited effective memory in recurrent designs. We present LoGeR (Long-context Geometric Reconstruction), a novel architecture that scales dense 3D reconstruction to extremely long sequences without post-optimization. LoGeR processes video streams in chunks, leveraging strong bidirectional priors for high-fidelity intra-chunk reasoning. To manage the critical challenge of coherence across chunk boundaries, we propose a learning-based hybrid memory module. This dual-component system combines a parametric Test-Time Training (TTT) memory to anchor the global coordinate frame and prevent scale drift, alongside a non-parametric Sliding Window Attention (SWA) mechanism to preserve uncompressed context for high-precision adjacent alignment. Remarkably, this memory architecture enables LoGeR to be trained on sequences of 128 frames, and generalize up to thousands of frames during inference. Evaluated across standard benchmarks and a newly repurposed VBR dataset with sequences of up to 19k frames, LoGeR substantially outperforms prior state-of-the-art feedforward methods--reducing ATE on KITTI by over 74%--and achieves robust, globally consistent reconstruction over unprecedented horizons.
title LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.03269