LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Junyi, Herrmann, Charles, Hur, Junhwa, Sun, Chen, Yang, Ming-Hsuan, Cole, Forrester, Darrell, Trevor, Sun, Deqing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
by: Zhang, Junyi, et al.
Published: (2024)
by: Zhang, Junyi, et al.
Published: (2024)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
by: Zhang, Junyi, et al.
Published: (2023)
by: Zhang, Junyi, et al.
Published: (2023)
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
by: Gu, Leslie, et al.
Published: (2025)
by: Gu, Leslie, et al.
Published: (2025)
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
by: Hur, Junhwa, et al.
Published: (2026)
by: Hur, Junhwa, et al.
Published: (2026)
Boundary Attention: Learning curves, corners, junctions and grouping
by: Polansky, Mia Gaia, et al.
Published: (2024)
by: Polansky, Mia Gaia, et al.
Published: (2024)
WonderJourney: Going from Anywhere to Everywhere
by: Yu, Hong-Xing, et al.
Published: (2023)
by: Yu, Hong-Xing, et al.
Published: (2023)
LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
Motion Prompting: Controlling Video Generation with Motion Trajectories
by: Geng, Daniel, et al.
Published: (2024)
by: Geng, Daniel, et al.
Published: (2024)
High-Resolution Frame Interpolation with Patch-based Cascaded Diffusion
by: Hur, Junhwa, et al.
Published: (2024)
by: Hur, Junhwa, et al.
Published: (2024)
DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
by: Xiong, Yajiao, et al.
Published: (2025)
by: Xiong, Yajiao, et al.
Published: (2025)
MASIV: Toward Material-Agnostic System Identification from Videos
by: Zhao, Yizhou, et al.
Published: (2025)
by: Zhao, Yizhou, et al.
Published: (2025)
A Simple Approach to Unifying Diffusion-based Conditional Generation
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
DreamWalk: Style Space Exploration using Diffusion Guidance
by: Shu, Michelle, et al.
Published: (2024)
by: Shu, Michelle, et al.
Published: (2024)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
Efficient Hybrid Zoom using Camera Fusion on Mobile Phones
by: Wu, Xiaotong, et al.
Published: (2024)
by: Wu, Xiaotong, et al.
Published: (2024)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
by: Nam, Jisu, et al.
Published: (2026)
by: Nam, Jisu, et al.
Published: (2026)
EA3D: Online Open-World 3D Object Extraction from Streaming Videos
by: Zhou, Xiaoyu, et al.
Published: (2025)
by: Zhou, Xiaoyu, et al.
Published: (2025)
DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
by: Zhou, Xiaoyu, et al.
Published: (2023)
by: Zhou, Xiaoyu, et al.
Published: (2023)
ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs
by: Shah, Viraj, et al.
Published: (2023)
by: Shah, Viraj, et al.
Published: (2023)
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
by: Gillman, Nate, et al.
Published: (2025)
by: Gillman, Nate, et al.
Published: (2025)
MotionV2V: Editing Motion in a Video
by: Burgert, Ryan, et al.
Published: (2025)
by: Burgert, Ryan, et al.
Published: (2025)
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
by: Feng, Haiwen, et al.
Published: (2025)
by: Feng, Haiwen, et al.
Published: (2025)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
Lumiere: A Space-Time Diffusion Model for Video Generation
by: Bar-Tal, Omer, et al.
Published: (2024)
by: Bar-Tal, Omer, et al.
Published: (2024)
Geometric Context Transformer for Streaming 3D Reconstruction
by: Chen, Lin-Zhuo, et al.
Published: (2026)
by: Chen, Lin-Zhuo, et al.
Published: (2026)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
by: Agrawal, Kumar Krishna, et al.
Published: (2025)
Selfi: Self Improving Reconstruction Engine via 3D Geometric Feature Alignment
by: Deng, Youming, et al.
Published: (2025)
by: Deng, Youming, et al.
Published: (2025)
Segment Anything without Supervision
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting
by: Zhou, Xiaoyu, et al.
Published: (2024)
by: Zhou, Xiaoyu, et al.
Published: (2024)
xT: Nested Tokenization for Larger Context in Large Images
by: Gupta, Ritwik, et al.
Published: (2024)
by: Gupta, Ritwik, et al.
Published: (2024)
LoBE-GS: Load-Balanced and Efficient 3D Gaussian Splatting for Large-Scale Scene Reconstruction
by: Hung, Sheng-Hsiang, et al.
Published: (2025)
by: Hung, Sheng-Hsiang, et al.
Published: (2025)
Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time Training
by: Liu, Changkun, et al.
Published: (2026)
by: Liu, Changkun, et al.
Published: (2026)
Emergent Temporal Correspondences from Video Diffusion Transformers
by: Nam, Jisu, et al.
Published: (2025)
by: Nam, Jisu, et al.
Published: (2025)
Gaga: Group Any Gaussians via 3D-aware Memory Bank
by: Lyu, Weijie, et al.
Published: (2024)
by: Lyu, Weijie, et al.
Published: (2024)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024)
by: Luo, Grace, et al.
Published: (2024)
GR3EN: Generative Relighting for 3D Environments
by: Xing, Xiaoyan, et al.
Published: (2026)
by: Xing, Xiaoyan, et al.
Published: (2026)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
Similar Items
-
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
by: Zhang, Junyi, et al.
Published: (2024) -
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
by: Zhang, Junyi, et al.
Published: (2023) -
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
by: Gu, Leslie, et al.
Published: (2025) -
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
by: Hur, Junhwa, et al.
Published: (2026) -
Boundary Attention: Learning curves, corners, junctions and grouping
by: Polansky, Mia Gaia, et al.
Published: (2024)