Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming Visual Geometry Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahdi, Soroush, Ayar, Fardin, Javanmardi, Ehsan, Tsukada, Manabu, Javanmardi, Mahdi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915530993041408
author Mahdi, Soroush
Ayar, Fardin
Javanmardi, Ehsan
Tsukada, Manabu
Javanmardi, Mahdi
author_facet Mahdi, Soroush
Ayar, Fardin
Javanmardi, Ehsan
Tsukada, Manabu
Javanmardi, Mahdi
contents Streaming visual transformers like StreamVGGT achieve strong 3D perception but suffer from unbounded growth of key value (KV) memory, which limits scalability. We propose a training-free, inference-time token eviction policy that bounds memory by discarding redundant tokens while keeping the most informative ones. Our method uses significantly less memory with little to no drop in accuracy: on 7-Scenes with long sequences it reduces peak memory from 18.63 GB to 9.39 GB while accuracy and completeness drop by only 0.003. Under strict memory budgets, eviction enables denser frame sampling, which improves reconstruction accuracy compared to the baseline. Experiments across video depth estimation (Sintel, KITTI), 3D reconstruction (7-Scenes, NRGBD), and camera pose estimation (Sintel, TUM-dynamics) show that our approach closely matches StreamVGGT at a fraction of the memory and makes long-horizon streaming inference more practical.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming Visual Geometry Transformers
Mahdi, Soroush
Ayar, Fardin
Javanmardi, Ehsan
Tsukada, Manabu
Javanmardi, Mahdi
Computer Vision and Pattern Recognition
Streaming visual transformers like StreamVGGT achieve strong 3D perception but suffer from unbounded growth of key value (KV) memory, which limits scalability. We propose a training-free, inference-time token eviction policy that bounds memory by discarding redundant tokens while keeping the most informative ones. Our method uses significantly less memory with little to no drop in accuracy: on 7-Scenes with long sequences it reduces peak memory from 18.63 GB to 9.39 GB while accuracy and completeness drop by only 0.003. Under strict memory budgets, eviction enables denser frame sampling, which improves reconstruction accuracy compared to the baseline. Experiments across video depth estimation (Sintel, KITTI), 3D reconstruction (7-Scenes, NRGBD), and camera pose estimation (Sintel, TUM-dynamics) show that our approach closely matches StreamVGGT at a fraction of the memory and makes long-horizon streaming inference more practical.
title Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming Visual Geometry Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.17650